Amazon Bedrock added Kimi K3, and its agent runtime now starts in about two seconds
Two Bedrock announcements landed on September 18, 2026. Kimi K3, which Moonshot AI describes as the first open model with 2.8 trillion parameters, is generally available, and it is the first open-weight model on Bedrock that supports explicit prompt caching. The runtime agents run inside was replaced at the same time: AWS measured a P75 cold start of 1.9 to 2.0 seconds for container images from 200 MB to 2 GB, against 5.4 to 30 seconds on the previous version. Memory is now released during a session instead of being held at the peak.
Source
Introducing Kimi K3 on Amazon BedrockAWS What's New — Machine Learning · Original published September 18, 2026
Amazon Bedrock published two changes on September 18, 2026, and they answer different questions. The first is which model you can call: Kimi K3 from Moonshot AI is now generally available, with a context window of 1 million tokens and native vision. The second is what your own agent runs inside: the next generation of AgentCore Runtime changes how memory is billed and how long a cold start takes.
What Kimi K3 brings, and who is claiming what
AWS states the availability and the platform behavior; the specifications of the model itself come from its maker. According to Moonshot AI, Kimi K3 is its most capable model and the first open model to reach 2.8 trillion parameters, with an improvement in scaling efficiency of approximately 2.5 times over Kimi K2. AWS adds that it combines native vision with a context window of 1 million tokens, and that is the part that changes the work: a large code repository or a stack of scanned pages fits into a single call.
The platform claim is AWS's own and is narrower: Kimi K3 is the first open-weight model on Amazon Bedrock to support explicit prompt caching. Bedrock also says your data is processed inside the AWS data boundary, is not shared with the model provider and is not used to train the model, that zero data retention is always enabled for inference requests, and that operators have no access during inference.
Prompt caching is where the cost changes
Explicit caching does not switch itself on. You mark the end of a reusable prompt prefix with a prompt_cache_breakpoint, after at least 1,024 tokens, and later requests that match that prefix reuse the stored content. Tokens written to the cache are billed at a higher rate, they stay cached for at least 30 minutes, and matching requests are billed at a discounted input rate.
There is a second lever, and it has nothing to do with caching. Kimi K3 is invoked through a cross-Region inference profile, and AWS says the global profile global.moonshotai.kimi-k3 costs approximately 10% less than a geographic one. The profile us.moonshotai.kimi-k3 keeps processing inside the United States for anyone bound by data residency rules. Availability is described as every AWS Region where Bedrock exists, through cross-Region inferencing.
AgentCore Runtime changed what you pay for and how long you wait
The second announcement replaces the serverless microVM compute that agents run inside on Amazon Bedrock. Each session now starts with a small memory profile, takes more on demand, and hands back memory that is no longer in use instead of holding it until the session ends. AWS puts it without hedging: you pay for actual usage rather than the peak.
Cold starts were rebuilt around a snapshot. The runtime prepares the agent environment once, snapshots it, and every new instance restores from that snapshot instead of repeating the startup sequence. In AWS's own testing that produced a P75 cold start of 1.9 to 2.0 seconds for container images from 200 MB to 2 GB, against 5.4 to 30 seconds on V1. The spread was the problem in itself, because a larger image meant a slower start.
- The new runtime is available in us-east-1, us-east-2, us-west-2, eu-west-1 and ap-northeast-1.
- It is switched on by hand: set
platformVersiontoV2when creating or updating a runtime. - Nothing else changes. There is no pre-provisioning, scaling still goes to zero, and session isolation is still enforced in hardware.
What this changes when you choose where to run an agent
Neither announcement is an independent measurement and neither should be read as one. The Kimi K3 figures are the description Moonshot AI gives of its own model, which AWS passes on. The cold-start figures are AWS measuring its own service against its own previous version. Both are still worth having, because these are the numbers the bill is calculated from rather than points on a leaderboard.
This suggests the interesting part is neither the model nor the runtime on its own, but that the two arrived on the same day. A context of 1 million tokens is expensive to resend on every call, and explicit caching and a runtime that releases idle memory are the two places where that cost comes down. One possible explanation for the timing is exactly that: long-context agents are the workload both changes were built for.
„you pay for actual usage rather than the peak“
Sources
Related
Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row
Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.
Microsoft, Hugging Faceverified

Ai2 releases AstaBrief, an open-weights model that writes cited research reports
Ai2 released the weights and training data for AstaBrief on October 2, 2026. The model takes a research question plus retrieved excerpts from the literature and writes a report with citations. Ai2 measures it at 51.1 seconds per report in its Asta platform, against 178.5 seconds for the Claude-powered mode beside it. The license is Apache 2.0, and Ai2 says its evaluation predates the current frontier models.
Ai2verified

Amazon's Strands Labs releases Decider 2B, an open decision model that runs locally
Strands Labs, the experimental arm of the Strands Agents project, released Strands Decider 2B on October 1, 2026. The model does not write text: it picks one of the options you give it, answers yes or no, or scores on a scale, and it attaches a confidence value to each answer. The team reports a median of 115 ms per decision on Nvidia's RTX 3090. Code, weights, training data, and scripts are public under Apache-2.0. If you build agents, this gives you a cheap local step for routing, tool checks, and guardrails.
Strands Agentsverified

