ACM CAIS 2026
Table of Contents
1. Overview
| Field | Value |
|---|---|
| Event | Inaugural ACM Conference on AI and Agentic Systems |
| Dates | (tutorials May 26; main conf May 27-29) |
| Location | DoubleTree by Hilton San Jose, 2050 Gateway Place, San Jose CA |
| URL | caisconf.org |
| Organizer | ACM |
| Program | 63 peer-reviewed papers, 46 system demos, workshops + tutorials |
| Scale | 219 submissions; ~28% acceptance; 115+ institutions; sold out |
| Framing | The systems around the model: compound AI as reliable software artifacts |
2. Notes
- Status: PAST (conference ran May 26-29, 2026). This is a notes/synthesis pass.
- The thesis: no venue existed for how AI systems are designed, optimized, evaluated, and maintained as reliable software artifacts – CAIS is that venue. It centers composition and orchestration over individual-model capability.
3. Keynotes
| Speaker | Affiliation | Angle |
|---|---|---|
| Andy Konwinski | Co-founder Databricks and Perplexity; Laude Institute | Open-source AI research funding, agent benchmarking |
| Thariq Shihipar | Anthropic – Claude Code (Member of Technical Staff) | Agentic coding tools; tool design tracks model ability |
| Percy Liang | Stanford; Together AI, Simile AI; creator of Marin | Foundation models, benchmarking rigor (HELM lineage) |
Shihipar's framing is the through-line: "the right tool design depends on the model's current abilities, and those abilities keep changing."
4. Focus areas
The official four pillars:
- Architectural patterns and composition – multi-component networks, verifier-based systems, RAG and tool-augmented designs.
- System optimization and efficiency – end-to-end pipeline optimization, cost-performance trade-offs, architecture search.
- Engineering and operations – MLOps for compound AI, monitoring, security in multi-component systems.
- Evaluation and benchmarking – metrics, reproducibility frameworks, comparative methodologies.
The Two Sigma preview breaks security and privacy out as a distinct fifth pillar (tool-execution threats, alignment for consequential systems).
5. Program structure
- 63 peer-reviewed research papers.
- 46 working system demonstrations (live implementations).
- Workshops and tutorials on May 26: agent skills, agentic software engineering, RL environments, discovery agents, healthcare AI. (Count TBC – the homepage says five, the program page says six.)
- A novel Operational Experience Reports track for production infrastructure accounts – explicitly valuing production knowledge alongside theoretical novelty (per Two Sigma: Lance Martin / Anthropic on Claude's managed agents; Raluca Popa / Google on Gemini security operations).
6. Notable research (attributed – Two Sigma preview)
- Tool interaction: "Do Agents Need to Plan Step-by-Step?" (Otani et al., Megagon); "OpaqueToolsBench" (Hallinan et al.); "XGrammar++" (Li et al.).
- Safety and robustness: "The Verifier Tax" (Sah et al.); "Willful Disobedience" (Sharma et al.); "Malice in Agentland" (Boisvert et al.).
- Efficiency: "Constant-Memory Retrieval via Koopman Operator Estimation" (Johansen and Sridhar); "AgentStop" (Pham et al.); "Robust Batch-Level Query Routing" (Markovic-Voronov et al.).
- Evaluation: "ViBench" / "Vibe Code Bench"; "Trace-Level Analysis of Information Contamination".
- Demos: Sherlock (error resolution), SkyDiscover (algorithmic discovery), Agent 4 (multi-agent coordination), Context Viewer (debugging).
7. Steering committee
Graham Neubig, Lingjiao Chen, Jeff Dean, Omar Khattab, Monica Lam, Thang Luong, Michele Catasta, Chris Potts, Naveen Rao, Dawn Song, Ion Stoica.
8. Sponsors
- Snorkel AI, Megagon, Databricks, Voaige, Snowflake, Two Sigma, OpenHands, Oracle, Buildkite, Cisco, Mithril, Bland AI.
9. Why it matters
CAIS is the field's bet that the durable engineering questions live in the "systems-around-the-model" layer, not in model weights – and that they deserve one venue with shared evaluation criteria rather than being scattered across ML, NLP, and systems tracks. The editorial hope (Two Sigma) is that CAIS produces the field's MapReduce papers: infrastructure accounts that guide a generation of practitioners.
10. Related
- Cross-Agent Collaboration – how concurrent agents coordinate (CAIS pillar 3/4)
- Multi-Agent Workflow Frameworks – composition/orchestration (pillar 1)
- Agent Memory Architectures – engineering/operations (pillar 3)
- BEAM Summit 2023 and Events index