AWS added model caching to SageMaker HyperPod and measured about 60% faster scale-out
SageMaker HyperPod can now keep model weights on a node's local NVMe and pre-pull the container image, so a pod that used to wait minutes for downloads starts in seconds. AWS reports benchmarks across models from 57 GB to 145 GB showing around 60% faster scale-out, and an image cache that removes over two minutes of pull time, a 97% reduction. Those figures are the vendor's own and arrived without a description of the test. On the same day AWS made TwelveLabs Marengo 3.0 available as an embedding model in Amazon Bedrock Managed Knowledge Base, so a knowledge base can index what a video shows rather than only what was said in it.
Source
Amazon SageMaker HyperPod now supports model caching for faster inference autoscaling and reduced cold startsAWS What's New — Machine Learning · Original published September 11, 2026
Cold start is the most expensive part of serving a large model behind an autoscaler. Before a new pod answers anything it pulls a container image from ECR and model weights from S3 or FSx, and on a 145 GB model that runs into tens of minutes. Amazon SageMaker HyperPod now has model caching, which takes both downloads out of the start-up path. AWS announced it on September 11, 2026, generally available in every region where HyperPod runs.
Two caches, and a fallback that makes them optional
Model caching has two independent parts. The weights cache stores model weights on the node's local NVMe, so a pod reads them from local storage instead of pulling them over the network from S3 or FSx. The image cache pre-pulls the container image onto cluster nodes, so a starting pod skips the ECR download entirely.
Either cache can miss. AWS says a pod that lands on a node without a warm cache falls back to pulling from the original source automatically, so a miss costs the old start-up time rather than a stuck or failed pod. You enable the feature through the HyperPod Inference Operator, by adding a modelCacheConfig section to an InferenceEndpointConfig or a JumpStartModel resource.
What the numbers say, and what they leave out
AWS reports benchmarks across models from 57 GB to 145 GB: around 60% faster scale-out, and a 97% reduction in image-pull time that removes over two minutes of waiting. It also says the benefit grows with model size, which follows from what the cache replaces, because a download scales with the weights and a local read does not.
The announcement does not describe the test. It names no models, no instance types, no baseline in seconds, and no method for telling a cold node from a warm one, so the figures are a vendor measurement rather than an independent one. Pricing is absent too, so the announcement does not say what holding weights on local NVMe costs.
Bedrock knowledge bases stop depending on the transcript
Amazon Bedrock Managed Knowledge Base could already search media by turning audio and video into text and embedding that text. TwelveLabs Marengo 3.0, available in the managed knowledge base from the same day, instead encodes visual scenes, speech and video cues directly into multimodal embeddings, which AWS describes as capturing meaning that transcription alone cannot.
The vectors are 512-dimensional, and a result carries the start and end time of the matching segment, so an application can jump to the moment instead of to the file. Segmentation is configurable. AWS calls the retrieval accuracy state of the art without naming a benchmark, and the announcement lists neither regions nor prices.
„Benchmarks across models from 57 GB to 145 GB show around 60% faster scale-out“
Sources
Related
Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row
Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.
Microsoft, Hugging Faceverified

Ai2 releases AstaBrief, an open-weights model that writes cited research reports
Ai2 released the weights and training data for AstaBrief on October 2, 2026. The model takes a research question plus retrieved excerpts from the literature and writes a report with citations. Ai2 measures it at 51.1 seconds per report in its Asta platform, against 178.5 seconds for the Claude-powered mode beside it. The license is Apache 2.0, and Ai2 says its evaluation predates the current frontier models.
Ai2verified

Amazon's Strands Labs releases Decider 2B, an open decision model that runs locally
Strands Labs, the experimental arm of the Strands Agents project, released Strands Decider 2B on October 1, 2026. The model does not write text: it picks one of the options you give it, answers yes or no, or scores on a scale, and it attaches a confidence value to each answer. The team reports a median of 115 ms per decision on Nvidia's RTX 3090. Code, weights, training data, and scripts are public under Apache-2.0. If you build agents, this gives you a cheap local step for routing, tool checks, and guardrails.
Strands Agentsverified
