Agentic AI Systems Can Plan and Act—But Safety Remains Unsolved
UC Berkeley's Dawn Song explains why autonomous AI agents represent a fundamental shift in capability and risk.

From chatbots to autonomous agents
Artificial intelligence systems are evolving beyond responding to prompts. The next generation of AI—called agentic AI—can reason through problems, select and use external tools, and execute multi-step tasks with minimal human supervision. Unlike large language models that generate text based on patterns, agentic systems can recognize when a problem requires a calculator and invoke one to ensure accuracy.
Dawn Song, a computer science professor at UC Berkeley and co-director of the Berkeley Center for Responsible, Decentralized Intelligence, outlined this shift in a recent interview with UC Berkeley News. The distinction matters: traditional chatbots provide information when asked, while agentic systems can pursue goals independently, coordinating across software tools, writing code, and adapting strategies as conditions change.
In August 2026, Berkeley hosted the Agentic AI Summit, drawing 5,000 in-person attendees and tens of thousands online. Researchers from OpenAI, Google, Amazon, and Meta joined academic teams to assess where the technology stands and what comes next.
Why it matters
Autonomous AI agents represent a category change in what machines can do without human intervention. That autonomy creates new attack surfaces and failure modes that existing safety frameworks weren't designed to address. Organizations deploying these systems will need to rethink oversight, testing, and accountability structures before agents operate in production environments.
Security challenges scale with autonomy
When AI systems can take actions rather than just generate text, the consequences of errors or manipulation expand significantly. Song highlighted prompt injection attacks, where adversaries craft inputs that override user instructions and redirect the agent toward malicious behavior. Mistakes can compound across long decision chains, and agents may drift from intended objectives without sufficient guardrails.
"Agentic AI changes the security landscape because these systems are no longer limited to generating information; instead, they can take actions in the world," Song explained. The systems may interact with software, access external tools, execute code, or coordinate with other agents—each interaction point introducing potential vulnerabilities.
Addressing these risks requires advances in evaluation methodologies, secure system architectures, improved interpretability, and stronger human oversight mechanisms. Song emphasized that safety and security research must evolve alongside capability improvements, not lag behind deployment.
Reliability trumps scale
Song identified three major technical obstacles facing agentic AI. First, reliability remains inconsistent—systems perform well on many tasks but struggle with long-horizon reasoning and recovering from unexpected situations. Second, evaluation frameworks haven't kept pace with autonomy; researchers need rigorous methods to measure whether systems are reliable, secure, and trustworthy under realistic conditions. Berkeley's AgentBeats platform represents one effort to establish standardized, reproducible agent evaluation.
Third, safety and security demand new approaches tailored to agentic systems' unique vulnerabilities.
Looking ahead, Song sees promise in combining AI with formal verification methods to generate software with machine-checkable correctness and security guarantees. "Ultimately, I believe the next major breakthroughs in AI won't come from scaling models alone," she said. "They'll come from making AI systems more reliable, more trustworthy and better aligned with human goals."
Democratizing access
Berkeley's Agentic AI massive open online course series has enrolled nearly 40,000 learners since launching in 2024, reaching students, practitioners, policymakers, and entrepreneurs worldwide. Song emphasized that AI's future will be shaped not only by research breakthroughs but by millions of people making decisions about design, deployment, and governance.
The details were first reported by UC Berkeley News following the August 2026 Agentic AI Summit.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

