AgentScout Logo Agent Scout

ArXiv cs.AI Weekly Papers — Week of June 4, 2026: Self-Evolving Agents and Multi-Agent Governance

31 papers collected this week with 25 agent-related papers (81%). Key trends: self-evolving agent frameworks surge (EvoDS, SkillPyramid, EvoDrive), LAP protocol fills agent-to-instrument gap, and domain benchmarks expose frontier model limitations.

AgentScout ·
#arxiv #ai-agents #papers #weekly-tracker #self-evolving-agents #multi-agent-systems
Analyzing Data Nodes...
SIG_CONF:CALCULATING
Verified Sources

Data Overview

  • Snapshot Week: 2026-05-28 to 2026-06-04
  • Tracker: ArXiv cs.AI Weekly Papers (view all snapshots: /tech/ai-agents/data/?tracker=arxiv-cs-ai-weekly)
  • Update Frequency: Weekly
  • Primary Sources: ArXiv cs.AI, ArXiv cs.CL

Key Facts

  • Who: 31 papers collected from ArXiv cs.AI and cs.CL categories
  • What: 25 agent-related papers (81%), including 12 multi-agent papers and 5 self-evolving agent frameworks
  • When: Week of May 28 - June 4, 2026
  • Impact: 3 new benchmarks, 1 new protocol (LAP), 7 papers with venue acceptance

🔺 Scout Intel: What Others Missed

Confidence: high | Novelty Score: 65/100

Three self-evolving agent papers (EvoDS, SkillPyramid, EvoDrive) appear in the same week, signaling a shift from static agent architectures toward autonomous skill acquisition. LAP protocol addresses a gap most coverage ignores: agent-to-instrument communication. While MCP handles model-to-tool and A2A handles agent-to-agent, LAP targets the physical instrument edge critical for autonomous scientific research. Hedge-Bench’s <16% frontier model performance on real hedge fund tasks exposes the gap between benchmark success and professional domain competence.

Key Implication: Agent frameworks are entering a consolidation phase where autonomous skill acquisition and standardized protocols replace manual prompt engineering. The 40% concentration on self-evolving systems suggests the field recognizes current limitations of static agent capabilities.

This Week’s Papers

#TitleArXiv IDTrendVenue/Improvement
1EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management2606.0384110KDD 2026, +28.9% over SOTA
2SkillPyramid: Hierarchical Skill Consolidation for Self-Evolving Agents2606.036929+38.0% reward, -27.7% steps
3LAP: Agent-to-Instrument Protocol for Autonomous Science2606.037559NEW protocol
4GAIATrace + Vidur-Agent: Multi-Model Agentic AI Systems Characterization2606.017258GAIATrace dataset, Vidur-Agent simulator
5Unified Context Evolution for LLM Agents2606.023048ALFWorld: 75.4% → 96.3%
6EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving2606.036788Self-improving LLM agents
7Hedge-Bench: Benchmarking Agents on Financial Reasoning Tasks2606.039187102 tasks, frontier <16%
8NovelAPIBench: Diagnosing Knowledge Gaps in LLM Tool Use2606.0365771.9K tasks, 5 domains
9Uncertainty-Aware Clarification with Information Gain2606.031357ICML 2026, +3.7% success rate
10Agentic CLEAR: Multi-Level Evaluation of LLM Agents2605.226087ACL

Self-Evolving Agent Frameworks

EvoDS (2606.03841) — Zherui Yang, Fan Liu, Yansong Ning, Hao Liu — KDD 2026

  • Focus: Autonomous data science with skill learning and adaptive context compression
  • Key Innovation: Self-evolving framework that acquires skills without manual intervention
  • Performance: +28.9% over SOTA on data science benchmarks

SkillPyramid (2606.03692) — Yuan Xiong et al.

  • Focus: Hierarchical skill consolidation for reusable experience
  • Key Innovation: Multi-level skill hierarchy enabling composition and reuse
  • Performance: +38.0% reward improvement, -27.7% steps on ALFWorld and WebShop

Unified Context Evolution (2606.02304) — Zixuan Zhu et al.

  • Focus: Gradient-free framework externalizing agent experience
  • Key Innovation: Typed Evolvable Context Units for memory management
  • Performance: ALFWorld 75.4% → 96.3%, WebShop 45.1% → 61.3%

EvoDrive (2606.03678) — Tong Nie et al.

  • Focus: Safety-critical autonomous driving scenario generation
  • Key Innovation: Pareto evolution via self-improving LLM agents
  • Domain: Autonomous driving

Multi-Agent Systems & Governance

LAP Protocol (2606.03755) — Linwu Zhu et al.

  • Type: Agent-to-Instrument Protocol
  • Gap Filled: Complements MCP (model-to-tool) and A2A (agent-to-agent)
  • Use Case: Autonomous scientific instruments

GAIATrace + Vidur-Agent (2606.01725) — Donghwan Kim et al.

  • Artifact: First token-level trace dataset for multi-model agentic systems
  • Tool: Vidur-Agent simulator for reproducible experiments
  • Benchmark: GAIA

Constraint State Governance (2605.10481) — Tianxiao Li

  • Focus: Safety in LLM multi-agent systems
  • Paradigm: Constraint drift prevention through state governance
  • Key Insight: Safe behavior must be maintained, not merely asserted

12 Angry AI Agents (2605.01986) — Ahmet Bahaddin Ersoz

  • Benchmark: Multi-agent decision-making using cinematic jury deliberation
  • Finding: 17/18 runs resulted in hung jury; anchoring is dominant failure mode
  • Insight: RLHF intensity determines deliberative flexibility

Benchmarks & Evaluation

BenchmarkDomainSizeKey Finding
Hedge-Bench (2606.03918)Financial reasoning102 tasksFrontier agents <16%
NovelAPIBench (2606.03657)Tool-use knowledge gaps1.9K tasks6 diagnostic categories
GAIATrace (2606.01725)Multi-agent tracesToken-levelFirst trace dataset
BigFinanceBench (2606.03829)Financial research workflows-Workflow-grounded

Protocols & Infrastructure

LAP (Agent-to-Instrument Protocol)

  • ArXiv: 2606.03755
  • Gap: Fills agent-to-instrument communication edge
  • Relation: Complements MCP (Anthropic) and A2A (Google)
  • Use Case: Autonomous scientific research

OpenAPI Documentation Agent-Ready

  • ArXiv: 2605.14312 — EASE 2026
  • Tool: Hermes multi-agent system
  • Result: Detected 2,450 smells in 600 endpoints
  • Purpose: MCP agent readiness

Continuum (KV Cache TTL)

  • ArXiv: 2511.02230
  • Focus: Multi-turn agent scheduling
  • Performance: 8x improvement in job completion time

Week-over-Week Summary

MetricThis WeekLast WeekChange
Total Papers315 (partial)+26
Agent-Related Papers255+20
Multi-Agent Papers121+11
Self-Evolving Agents50NEW
Avg Trend Score (Agent)6.47.2-0.8
Accepted Papers (venue)71+6

Notable Additions This Week:

  • EvoDS (KDD 2026) — first self-evolving data science agent with accepted venue
  • LAP protocol — new protocol category (agent-to-instrument)
  • Hedge-Bench — exposes frontier model gap in professional tasks
  • SkillPyramid — hierarchical skill consolidation framework

Papers from Last Week (Now Ranked Lower):

  • MUSE-Autoskill (2605.27366) — Trend: 8 → N/A
  • SIA (2605.27276) — Trend: 8 → N/A
  • FinHarness (2605.27333) — Trend: 7 → N/A
  • QUACK (2605.27068) — Trend: 7 → N/A
  • Alignment Tampering (2605.27355) — Trend: 6 → N/A
  1. Self-evolving agent frameworks surge: 3 major papers (EvoDS, SkillPyramid, EvoDrive) focus on autonomous skill acquisition, representing 40% of top-10 papers by trend score

  2. Multi-agent governance emerging: LAP protocol fills agent-to-instrument gap, Constraint State Governance addresses safety in LLM multi-agent systems

  3. Domain-specific benchmarks proliferate: Hedge-Bench (finance), NovelAPIBench (tool-use), BigFinanceBench reveal specialized evaluation needs

  4. Context management critical: Unified Context Evolution demonstrates 96.3% on ALFWorld through typed Evolvable Context Units

  5. Multi-agent characterization tools: GAIATrace + Vidur-Agent enable reproducible simulation of multi-model agentic systems

  6. RLHF alignment intensity key: 12 Angry AI Agents shows alignment level determines deliberative flexibility in multi-agent settings

Category Distribution

CategoryCountPercentage
cs.AI1858%
cs.CL413%
cs.MA413%
cs.SE26%
cs.DC13%
cs.OS13%
Other13%

Accepted Papers (with Venue)

PaperVenueArXiv ID
EvoDSKDD 20262606.03841
Uncertainty-Aware ClarificationICML 20262606.03135
Agentic CLEARACL2605.22608
Cattle TradeICLR 2026 Workshop2605.14537
OpenAPI DocumentationEASE 20262605.14312
LLM Agent SystemsIEEE AIIoT 20252505.16120
When to Re-PlanICML 2026 Workshop2606.03741

Previous Snapshots

This is the first snapshot for the ArXiv cs.AI Weekly Tracker. Future snapshots will be linked here.


Sources


Last updated: 2026-06-04 by AgentScout automated tracker. Collection duration: 180 seconds. Sources: 2/4 succeeded (ArXiv direct API rate-limited, HuggingFace 404).

ArXiv cs.AI Weekly Papers — Week of June 4, 2026: Self-Evolving Agents and Multi-Agent Governance

31 papers collected this week with 25 agent-related papers (81%). Key trends: self-evolving agent frameworks surge (EvoDS, SkillPyramid, EvoDrive), LAP protocol fills agent-to-instrument gap, and domain benchmarks expose frontier model limitations.

AgentScout ·
#arxiv #ai-agents #papers #weekly-tracker #self-evolving-agents #multi-agent-systems
Analyzing Data Nodes...
SIG_CONF:CALCULATING
Verified Sources

Data Overview

  • Snapshot Week: 2026-05-28 to 2026-06-04
  • Tracker: ArXiv cs.AI Weekly Papers (view all snapshots: /tech/ai-agents/data/?tracker=arxiv-cs-ai-weekly)
  • Update Frequency: Weekly
  • Primary Sources: ArXiv cs.AI, ArXiv cs.CL

Key Facts

  • Who: 31 papers collected from ArXiv cs.AI and cs.CL categories
  • What: 25 agent-related papers (81%), including 12 multi-agent papers and 5 self-evolving agent frameworks
  • When: Week of May 28 - June 4, 2026
  • Impact: 3 new benchmarks, 1 new protocol (LAP), 7 papers with venue acceptance

🔺 Scout Intel: What Others Missed

Confidence: high | Novelty Score: 65/100

Three self-evolving agent papers (EvoDS, SkillPyramid, EvoDrive) appear in the same week, signaling a shift from static agent architectures toward autonomous skill acquisition. LAP protocol addresses a gap most coverage ignores: agent-to-instrument communication. While MCP handles model-to-tool and A2A handles agent-to-agent, LAP targets the physical instrument edge critical for autonomous scientific research. Hedge-Bench’s <16% frontier model performance on real hedge fund tasks exposes the gap between benchmark success and professional domain competence.

Key Implication: Agent frameworks are entering a consolidation phase where autonomous skill acquisition and standardized protocols replace manual prompt engineering. The 40% concentration on self-evolving systems suggests the field recognizes current limitations of static agent capabilities.

This Week’s Papers

#TitleArXiv IDTrendVenue/Improvement
1EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management2606.0384110KDD 2026, +28.9% over SOTA
2SkillPyramid: Hierarchical Skill Consolidation for Self-Evolving Agents2606.036929+38.0% reward, -27.7% steps
3LAP: Agent-to-Instrument Protocol for Autonomous Science2606.037559NEW protocol
4GAIATrace + Vidur-Agent: Multi-Model Agentic AI Systems Characterization2606.017258GAIATrace dataset, Vidur-Agent simulator
5Unified Context Evolution for LLM Agents2606.023048ALFWorld: 75.4% → 96.3%
6EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving2606.036788Self-improving LLM agents
7Hedge-Bench: Benchmarking Agents on Financial Reasoning Tasks2606.039187102 tasks, frontier <16%
8NovelAPIBench: Diagnosing Knowledge Gaps in LLM Tool Use2606.0365771.9K tasks, 5 domains
9Uncertainty-Aware Clarification with Information Gain2606.031357ICML 2026, +3.7% success rate
10Agentic CLEAR: Multi-Level Evaluation of LLM Agents2605.226087ACL

Self-Evolving Agent Frameworks

EvoDS (2606.03841) — Zherui Yang, Fan Liu, Yansong Ning, Hao Liu — KDD 2026

  • Focus: Autonomous data science with skill learning and adaptive context compression
  • Key Innovation: Self-evolving framework that acquires skills without manual intervention
  • Performance: +28.9% over SOTA on data science benchmarks

SkillPyramid (2606.03692) — Yuan Xiong et al.

  • Focus: Hierarchical skill consolidation for reusable experience
  • Key Innovation: Multi-level skill hierarchy enabling composition and reuse
  • Performance: +38.0% reward improvement, -27.7% steps on ALFWorld and WebShop

Unified Context Evolution (2606.02304) — Zixuan Zhu et al.

  • Focus: Gradient-free framework externalizing agent experience
  • Key Innovation: Typed Evolvable Context Units for memory management
  • Performance: ALFWorld 75.4% → 96.3%, WebShop 45.1% → 61.3%

EvoDrive (2606.03678) — Tong Nie et al.

  • Focus: Safety-critical autonomous driving scenario generation
  • Key Innovation: Pareto evolution via self-improving LLM agents
  • Domain: Autonomous driving

Multi-Agent Systems & Governance

LAP Protocol (2606.03755) — Linwu Zhu et al.

  • Type: Agent-to-Instrument Protocol
  • Gap Filled: Complements MCP (model-to-tool) and A2A (agent-to-agent)
  • Use Case: Autonomous scientific instruments

GAIATrace + Vidur-Agent (2606.01725) — Donghwan Kim et al.

  • Artifact: First token-level trace dataset for multi-model agentic systems
  • Tool: Vidur-Agent simulator for reproducible experiments
  • Benchmark: GAIA

Constraint State Governance (2605.10481) — Tianxiao Li

  • Focus: Safety in LLM multi-agent systems
  • Paradigm: Constraint drift prevention through state governance
  • Key Insight: Safe behavior must be maintained, not merely asserted

12 Angry AI Agents (2605.01986) — Ahmet Bahaddin Ersoz

  • Benchmark: Multi-agent decision-making using cinematic jury deliberation
  • Finding: 17/18 runs resulted in hung jury; anchoring is dominant failure mode
  • Insight: RLHF intensity determines deliberative flexibility

Benchmarks & Evaluation

BenchmarkDomainSizeKey Finding
Hedge-Bench (2606.03918)Financial reasoning102 tasksFrontier agents <16%
NovelAPIBench (2606.03657)Tool-use knowledge gaps1.9K tasks6 diagnostic categories
GAIATrace (2606.01725)Multi-agent tracesToken-levelFirst trace dataset
BigFinanceBench (2606.03829)Financial research workflows-Workflow-grounded

Protocols & Infrastructure

LAP (Agent-to-Instrument Protocol)

  • ArXiv: 2606.03755
  • Gap: Fills agent-to-instrument communication edge
  • Relation: Complements MCP (Anthropic) and A2A (Google)
  • Use Case: Autonomous scientific research

OpenAPI Documentation Agent-Ready

  • ArXiv: 2605.14312 — EASE 2026
  • Tool: Hermes multi-agent system
  • Result: Detected 2,450 smells in 600 endpoints
  • Purpose: MCP agent readiness

Continuum (KV Cache TTL)

  • ArXiv: 2511.02230
  • Focus: Multi-turn agent scheduling
  • Performance: 8x improvement in job completion time

Week-over-Week Summary

MetricThis WeekLast WeekChange
Total Papers315 (partial)+26
Agent-Related Papers255+20
Multi-Agent Papers121+11
Self-Evolving Agents50NEW
Avg Trend Score (Agent)6.47.2-0.8
Accepted Papers (venue)71+6

Notable Additions This Week:

  • EvoDS (KDD 2026) — first self-evolving data science agent with accepted venue
  • LAP protocol — new protocol category (agent-to-instrument)
  • Hedge-Bench — exposes frontier model gap in professional tasks
  • SkillPyramid — hierarchical skill consolidation framework

Papers from Last Week (Now Ranked Lower):

  • MUSE-Autoskill (2605.27366) — Trend: 8 → N/A
  • SIA (2605.27276) — Trend: 8 → N/A
  • FinHarness (2605.27333) — Trend: 7 → N/A
  • QUACK (2605.27068) — Trend: 7 → N/A
  • Alignment Tampering (2605.27355) — Trend: 6 → N/A
  1. Self-evolving agent frameworks surge: 3 major papers (EvoDS, SkillPyramid, EvoDrive) focus on autonomous skill acquisition, representing 40% of top-10 papers by trend score

  2. Multi-agent governance emerging: LAP protocol fills agent-to-instrument gap, Constraint State Governance addresses safety in LLM multi-agent systems

  3. Domain-specific benchmarks proliferate: Hedge-Bench (finance), NovelAPIBench (tool-use), BigFinanceBench reveal specialized evaluation needs

  4. Context management critical: Unified Context Evolution demonstrates 96.3% on ALFWorld through typed Evolvable Context Units

  5. Multi-agent characterization tools: GAIATrace + Vidur-Agent enable reproducible simulation of multi-model agentic systems

  6. RLHF alignment intensity key: 12 Angry AI Agents shows alignment level determines deliberative flexibility in multi-agent settings

Category Distribution

CategoryCountPercentage
cs.AI1858%
cs.CL413%
cs.MA413%
cs.SE26%
cs.DC13%
cs.OS13%
Other13%

Accepted Papers (with Venue)

PaperVenueArXiv ID
EvoDSKDD 20262606.03841
Uncertainty-Aware ClarificationICML 20262606.03135
Agentic CLEARACL2605.22608
Cattle TradeICLR 2026 Workshop2605.14537
OpenAPI DocumentationEASE 20262605.14312
LLM Agent SystemsIEEE AIIoT 20252505.16120
When to Re-PlanICML 2026 Workshop2606.03741

Previous Snapshots

This is the first snapshot for the ArXiv cs.AI Weekly Tracker. Future snapshots will be linked here.


Sources


Last updated: 2026-06-04 by AgentScout automated tracker. Collection duration: 180 seconds. Sources: 2/4 succeeded (ArXiv direct API rate-limited, HuggingFace 404).

bdxu9juti5hm47y5idvql░░░ya9obd3pkan8p4u47jcqt49edlbql2brt████iizkldggcys0o7na96vzc2dhyqb50ad0je░░░r4xjcroyob9uvyrm6gs50e6jio3tr9ybi░░░vcka8oiaph2h7ek18bp8qqpydlsq7z1f░░░uucm1e7fw52eth18osb7dv173uuqlsoo████ogr6tlrwmi412mwkx6i1e34d8nlom4d████hcmmr9l89lvm0z3xvm63q1u4r6nmwflo░░░i68kvw6b3ov5k3m8sjm5anbn4ryy7iu░░░2wztig9oupteuzf0ksjo0679cgh7wba24████b5uhdut452xlk1h06m9m79vum4yi0pxj████u3utyofagd9q845z0ald48pedgc2ukwye████xqqe5mvpreaijte4swsouu0k8z5n2p3████42oiv34gnrk3h0euam3m6phdmqlq18c8░░░vovwukzji9k3se5a188vyssrgdzwnh2mf░░░n45ce7uy2rh5py2reuumm1m41ie2uifv░░░cd7y27jmtagz7a2r0ld4874njhdfwkbu████zxcgtbxpsbh39m00w9x4l6yiunkyeq9░░░83xn20yyd7v2ybdy6cf8mnky7p1lr6n████gb50vvbw9xn17yp6e8gjilc5v13jhj1f░░░g6mgtrk45web15t69e8fbawy965ewd87████nyofposzq0tftc6mu14yp74emm84b5zjq░░░11tzq90feq3y6wfszfc0fk4fhuiuswxug░░░vwemmy4hxefl8q3l53ko787atzpusqwo████7uv2kxgljg6ft8pmqtskhyp95t4xbv5k████tetgtn2p8dgrnneemnjf75c2xh252bqn████qt2gjwk5qrlx5ft7llbsvqacjerubb7n9████fdqs74ppzw8vrdai9t2pi9ff4berwipuj████xh04tfak0gffoihzvocajuxxrns4zuqxd░░░g5a9h9l3rkwqm8ntriw1jf4lhs4cdzlo░░░z07eh7nn8ikfk9eaueitb84p3i8jvji░░░gwlxwrxzub7hkndehb8nvaudynukf95qg████p460058ujfzo2iqygnp1gvs0tsci6tub░░░44umhbx72opjiuyrucpqmkh6djom7lvlw░░░ia2xgx40jtibc5gphvzun8v62xa77ikp████cm8qq8ca9q7vps2700snc8mluw93qdms████sz6egdvbt4dpfqyp2jdmb9dh2igwcb3j4░░░wpi9i8b7g9av1bc0sa3pczl7lc830mff░░░i7c3yj8ma3bjgl54ztjwht027fmu1nek3q░░░tm7tpwzq40n4ak9pe3b56noz3hxmcyzli████j23tp7lngdmchouxy293zsaps2tjkfzuq░░░djb5bxrqtlff0he2sqra8famc7xdp12ci████cewt5xbdksu7xnej5h8lahdc6kqul8zc████kvf7wjldr5lpmua30xzmnddz1ksqfsc6████fofwnmhcztkk318u2ytm0bqhcq96a5cb████9cq4dlctoz2lusdnp14jtml7428cvv99████tnmsxpqpto4vegfmvpqjgswh5jiswq59████e5nsmvhz6e8bv5y4o69s4tuqfmotacs3m████tg7va6bkfl2nuds9onxp59oa16n37si░░░zrdme2ctm3jkkyzy7dz02h8bj5o4lw5████olkwwpuefe