AWS expands Bedrock with million-token context, agent governance
August updates bring cross-region inference, 14-day agent sessions, spending controls, and robotics integration to Amazon's AI platform.
Amazon Web Services shipped a wave of enterprise AI capabilities in August 2026, extending Amazon Bedrock's foundation model platform with tools designed for production workloads that span longer time horizons, stricter governance requirements, and physical environments.
The updates address a shift in how organizations are deploying AI: moving from isolated model calls to systems that handle multi-day workflows, operate within defined spending and security boundaries, and coordinate actions across digital and physical infrastructure.
Million-token context and global routing
AWS added million-token context windows with prompt caching to OpenAI's GPT-5.6 Sol, Terra, and Luna models on Bedrock, enabling applications to process entire codebases or regulatory documents in a single request. The models now include integrated web search with citation support, eliminating the need for separate search provider integrations.
Cross-region inference expanded to more than 25 AWS Regions, allowing developers to route requests globally while maintaining geographic data-processing boundaries through Geo profiles. OpenAI simultaneously reduced pricing across the GPT-5.6 family on Bedrock.
New cost management features include IAM principal cost allocation, which attributes inference spending to specific users, teams, or projects, and AWS Cost Anomaly Detection support for third-party foundation models.
Agent runtime and governance
Amazon Bedrock AgentCore introduced runtime instances that support agent sessions lasting up to 14 days on dedicated EC2 compute, including GPU-accelerated and memory-optimized configurations. This infrastructure enables agents to handle research, coding, and monitoring tasks that unfold over multiple days.
Temporal policies enforce action sequences, prerequisites, and approval gates based on an agent's prior activity. Rate limiting controls requests, inference tokens, and concurrent connections by user or group. A new payments capability allows agents to access paid APIs and content with infrastructure-enforced spending limits and full observability.
AgentCore memory expanded beyond conversational data to extract long-term context from activity logs, behavioral events, and structured JSON sources, with fine-grained access control that isolates memories by user or tenant. Web search for agents includes domain filtering and publication date controls.
The AWS Agent Registry provides a searchable catalog for agents, MCP servers, and skills across connected AWS accounts, with organization-wide detection and integration into Amazon Quick.
Regulated environments and physical systems
Claude Opus 5 and OpenAI GPT-5.6 Terra and Luna became available in AWS GovCloud (US) Regions, with zero data retention enabled by default for Claude. Amazon Nova Multimodal Embeddings and AgentCore capabilities also expanded to GovCloud.
For robotics applications, Strands Robots connects Strands Agents, LeRobot, and Hugging Face Storage Buckets into a unified workflow for recording demonstrations, training policies, and deploying to simulated or physical hardware. The platform supports mesh-based discovery and coordination across devices using Zenoh for local networks and AWS IoT Core for distributed fleets.
Why it matters
These updates reflect a maturation in enterprise AI deployment beyond proof-of-concept demos. Organizations building production AI systems need infrastructure that can process large context windows economically, maintain agent state across extended workflows, enforce spending and security boundaries programmatically, and extend into physical operations. AWS is positioning Bedrock and AgentCore as a platform for AI systems that handle meaningful work autonomously while operating within organizational controls—a requirement for regulated industries and high-stakes applications where model output alone isn't sufficient.
Details were first reported by AWS in a September 9, 2026 blog post by Tanvi Girinath.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
