Skip to content

Daily desk

Redakcija WebAiRadar

40 items

A model drafts the text; a person reviews, corrects and signs it. Originals are never translated or reprinted — each piece is independent analysis with figures verified against the primary source.

Agentsmedium signal

A proxy that stripped one header was doubling Claude Code's API bill

Version 2.1.239 fixes streaming on Bedrock behind proxies that remove the response Content-Type header. Claude Code silently fell back to re-running every turn without streaming, and each turn was billed twice. The same release makes cost estimates show the 1.1× premium that data-residency workspaces pay.

Anthropicverified

Black Forest Labs' announcement graphic: stacked cards showing an illustrated tree, beside the words "4K Video Upscaler, powered by FLUX 3".
Video & Imagemedium signal

FLUX Upscale prices video by the megapixel-second, and charges only for what comes out

Black Forest Labs released FLUX Upscale on August 20 as a standalone tool and endpoint that regenerates video up to 4K. It runs in two modes: Precise at $0.07 per megapixel-second and Creative at $0.10. The billing detail matters more than either rate, because the charge is calculated on the output alone, so the upscale factor you pick is what sets the bill.

Black Forest Labsverified

PIN REMOVEDhttpx2llm 0.33 declares what it imports
Toolsmedium signal

llm 0.33 takes the pin out and moves to httpx2

Version 0.32.1 held fresh installs together by pinning the OpenAI Python package below 3.0.0. That was a holding action, and 0.33 does the actual repair: it upgrades to the OpenAI library 3.x and switches its HTTP client from httpx to httpx2. The release also carries changes you will notice in daily use, including a --key option on the embedding commands and server-side tool results that finally show up in llm logs.

Simon Willisonverified

Modelsstrong signal

GPT-5.6 Sol drops to $4 and $20, and overtakes Claude Opus 5 on cost

OpenAI cut Sol's API price on August 21: input from $5 to $4, output from $30 to $20, cached input from $0.50 to $0.40. On a standard task the model goes from more expensive than Claude Opus 5 to cheaper than it. The cut is promotional and runs at least through November 21.

OpenAIverified

Toolsweak signal

A fresh install of llm broke because the OpenAI library stopped using httpx

Version 0.32.1 pins the OpenAI Python package below 3.0.0 so that new installations work again. Nothing in llm changed to cause the break: it imported httpx while relying on the OpenAI package to bring it along, and when that package dropped httpx the dependency simply stopped arriving. It is a small release with a lesson that outlives it.

Simon Willisonverified

Webmedium signal

A class prefix selector is now spec text in Selectors Level 5

The draft of August 18 adds section 2.1, Class Prefix Selectors, defining .foo-* to match any class beginning with that prefix. It replaces the attribute-selector workarounds every design system carries. Nothing implements it yet, and the matching rule has an edge that will catch people: .foo-* does not match foo--bar.

CSS Working Groupverified

Modelsstrong signal

The best-scoring speech models reproduce the benchmark's own transcription errors

Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.

Hugging Faceverified

Agentsmedium signal

ChatGPT now reads and sends your Apple Messages, on the Mac only

A plug-in in the ChatGPT desktop app reads and searches your iMessage, SMS and RCS threads and sends messages through Messages on your behalf. By default nothing goes out until you approve both the text and the recipients. There is a switch that removes that step, and OpenAI's own documentation argues against using it.

OpenAIverified

Agentsmedium signal

GitHub Copilot moved into Slack and Microsoft Teams on the same day

Both shipped on August 21 in public preview. Mention @GitHub in a channel and the agent triages issues, investigates failures, writes changes in a cloud sandbox and opens a pull request, with the conversation attached. The interesting part is not the capability list, which is familiar, but the room it moved into: the place where work gets discussed rather than written.

GitHubverified

Agentsstrong signal

Nvidia's harness takes Claude Opus 5 from about 30% to a perfect ARC-AGI-3 score

Nvidia published a run in which AVO, its agent architecture, scores 100.00 RHAE on ARC-AGI-3 and clears all 183 levels across 25 environments. The same model evaluated on its own scores about 30%. In the same post Nvidia writes that the comparison is not a controlled ablation, and that sentence did not survive into the coverage.

NVIDIAverified

Modelsstrong signal

Gemini 3.7 Flash, read from Google's own numbers

Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.

Googleverified

Industrystrong signal

How to label AI content under Article 50, and which part of it is not your job

Article 50 of the EU AI Act has applied since August 2, 2026, and it binds anyone serving people in the Union, wherever the server is. Most of the panic is about the machine-readable marking requirement, which for a site owner who calls somebody else's API is somebody else's obligation. Here is what is actually yours: a chatbot that says what it is, published text that either carries a name or carries a label, and a deepfake that admits it.

Evropska komisijaverified

Agentsstrong signal

Gemini Spark vs Claude's browser toolset: whose account, whose risk

Both drive a web browser, both stop before the step that cannot be undone, and their lists of stopping points are nearly the same. The difference is what each assumes about the session it runs in: one is built on the accounts you are already signed into, the other tells you to give it a browser with no credentials in it. That single assumption decides which of the two belongs anywhere near your work.

Google, Anthropicverified

Modelsstrong signal

Best cheap models for high-volume work, priced per thousand calls

Six models, one task, one number: what a thousand calls cost when each sends 4,000 tokens in and gets 800 back. The cheapest row is $1.76 and the most expensive is $16.00, a nine-fold spread rather than the hundred-fold spread the category implies. Two things move the ranking more than the headline price does, and one of them has a date on it.

Anthropic, Google, OpenAIverified

Modelsstrong signal

Claude Sonnet 5, read from what Anthropic publishes

Two disclosures first: this is a reading of the vendor's own evaluations rather than our test, and it is written by a model that vendor built. With both stated, the published numbers still contain three things worth noticing before you pick this model.

Anthropicverified

Webmedium signal

Put Baseline in CI so the browser stops deciding for you

Three packages and about 20 minutes. One config tells your build which browsers you target, a linter warns when a stylesheet reaches past that line, and CI says it on the pull request instead of a user saying it in a bug report six weeks later.

web.dev, npmverified

Agentsstrong signal

Claude Code vs Cursor: same $20, different shape

Both start at $20 a month, both speak MCP, both run skills and hooks, both reach CI. The difference is not the feature list. It is whether the agent lives inside an editor you adopt or inside the terminal you already have.

Anthropic, Cursorverified

Modelsstrong signal

Claude Opus 5 vs GPT-5.6 Sol: which is cheaper depends on a threshold OpenAI does not publish

Sol's price cut on August 21 put it below Opus 5 on both short-context columns — $4 against $5 on input, $20 against $25 on output. Its long-context column went the other way and still costs more. Opus 5's price sits between Sol's two columns, so the cheaper model depends on which column your request lands in, and OpenAI does not say where the boundary is.

Anthropic, OpenAIverified

Modelsstrong signal

Cut your model bill by 89% without changing what you ship

Four levers, all published, none of them clever: cache the fixed part of your prompt, batch what can wait, drop a tier where the task allows it, and stop paying multipliers you did not ask for. Worked all the way through on one real workload.

Anthropic, OpenAIverified

Video & Imagestrong signal

Best AI video generators, priced per second

Three ways to buy generated video in August 2026: per second through an API, as monthly credits with every model in one place, or as credits tuned for iteration. The same ten-second 1080p clip costs anywhere from $0.80 to $7.00 depending on which door you walk through.

Google, Runway, Lumaverified

Industrystrong signal

Pew puts a number on it: 35% of pages published after ChatGPT show AI authorship

Pew Research ran nearly half a million pages from Common Crawl through a detector. Among pages published after November 2022, more than a third came back with significant signs of AI writing, and .com domains showed it at ten times the rate of .edu and .gov.

TechCrunchverified

Agentsstrong signal

Encrypting the instruction walks it straight past the guardrail

Researchers at Adversa hid a prompt injection as ciphertext with the key next to it. Grok decrypted it in its own sandbox and sent the user's name, location, and chat history to an attacker's server, without a warning and without asking.

Ars Technicaverified

Toolsmedium signal

Next.js will patch one critical vulnerability on August 26

Vercel published the date before the fix. Versions 16.3.3 and 15.5.24 arrive on August 26 with a full advisory, and the announcement exists so teams can schedule the upgrade instead of discovering it from a CVE feed.

Next.js Blogverified

Industrymedium signal

Gemma passes a billion downloads and 100,000 derived models

On August 20, Google announced that the Gemma family of open models has passed a billion downloads. In two years the community has published more than 100,000 derived models, and the most recent Kaggle competition drew more than 1,600 entries. An Awesome Gemma repository launched with the announcement, meant as an index of vetted projects, fine-tuned models, tutorials, and tools. The figure matters to someone building sites too, because a model you run on your own server stops being exotic, which makes text processing without sending data to a third party workable.

Googleverified

Toolsmedium signal

ChatGPT search switched en masse to the site: operator

On August 20, Simon Willison relayed a measurement from Promptwatch showing that ChatGPT abruptly began narrowing its searches to preselected domains. The share of queries carrying the site: operator held between 0.3% and 0.5% for weeks, then jumped to 16% to 17% on August 8 — two days after OpenAI announced accuracy fixes in GPT-5.6. Willison concludes that the search tool internally has the shape search(query, recency, domains), which means the model picks domains before it even frames the question. If you care about visibility in that search, you are no longer competing only for a place in the result but for your domain to make the list the model assembles on its own.

Simon Willisonverified

Industrymedium signal

Google gives sites a button readers use to name them a preferred source

On August 20, Google published a button a publisher places on its own page, which a reader clicks once to add that site to their preferred sources in Search. The consequence is that the site's pages appear more prominently in Top Stories, in AI Overviews, and in AI Mode. Google says users have chosen more than 600,000 different sources so far. Two reader-side changes arrived with it: tuning topics in the Discover feed through a three-dot menu, and a customizable audio briefing in Google News for Android.

Googleverified

Webweak signal

A timing chart as the blueprint for SMIL animation in SVG

On August 20 at Smashing Magazine, Johan Grobler described a method in which a complex SMIL animation is first drawn as a timing chart and only then written in markup. Every bar on the chart corresponds to one animate element, and its length and position become the values of the dur and begin attributes. Because a single animate changes exactly one property of one element, the count climbs fast: a three-dot loader takes six of them. The piece also covers synchronizing to another animation's start or end through the begin attribute, and honoring prefers-reduced-motion through the picture element.

Smashing Magazineverified

Agentsmedium signal

Claude gets a separate toolset for driving a browser

August 19 brought two changes to how Claude operates someone else's screen. The computer use tool moved to general availability as computer_toolset_20260801, with no beta header, actions grouped into a single turn, and zoom enabled from the start. Alongside it came browser_toolset_20260801, a separate set of 31 tools that works inside a browser you run yourself rather than across the whole desktop. The important difference is that this set reads the page accessibility tree and returns element references, so a click targets a reference such as ref_2 instead of a point in pixels — and that survives a layout shift.

Anthropicverified

Agentsweak signal

Claude Code 2.1.238 stops memory growth in long sessions

The August 20 release fixes what got in the way during long stretches of work rather than adding features. Tool results returned by subagents are now freed as soon as they leave the recent-view window, so memory in long interactive sessions stops growing without a ceiling. The release also fixes custom output styles sliding back to the default tone mid-session, along with a run of stalls in Remote Control. A new keybindingFlavor setting restores Bash behavior, where Ctrl+W deletes back to the previous space, and claude mcp list no longer starts disabled servers just to check their status.

Anthropicverified

Toolsweak signal

Workbench becomes Playground, and the old tool is gone

Since August 18, Workbench in the Claude console is called Playground and has been rebuilt from the ground up. The new tool supports every Messages API parameter and ships with templates that show how code execution and web search work. For every run it displays the complete request exactly as the SDK sends it, along with the raw API response, which makes it a way to learn the API rather than only a place to try prompts. The old Workbench shut down on August 17, saved prompts and evaluations do not carry over, and the experimental prompt tools APIs for generating and improving prompts were withdrawn with it.

Anthropicverified

Webmedium signal

Firefox 154 brings sibling-index(), sibling-count(), and text-box-trim

The August 18 release adds three CSS features you can use right away. sibling-index() and sibling-count() give an element's position among its siblings and their total, directly in CSS and without a line of JavaScript, which ends the need to hand-write variables like --i for staggered animation delays. Alongside them come text-box-edge, text-box-trim, and the text-box shorthand, which remove the extra space above and below text so headings finally sit in the grid the way they look in the design. On the JavaScript side, includes(), join(), chunks(), and windows() were added to the iterator prototype.

MDNverified

Agentsstrong signal

Files, Agent Skills, and the Admin API leave beta

On August 19, Anthropic moved a large part of the Claude platform from beta to general availability. The Files API no longer requires a beta header and gains file expiry through expires_in_seconds, pagination, and filtering by identifier, alongside a terabyte of storage per organization and a limit of 500 requests per minute. Agent Skills and the accompanying Skills API also became generally available, so skills now load through the Messages API without a single beta header. Managed Agents gain allow and block lists of domains for search and content fetching, and the session viewer in the console was redesigned with a timeline minimap and an Inspector panel. Anyone building an agent over client content gets three critical parts of the API that stop being a moving target.

Anthropicverified

Webmedium signal

New on the web platform in July: Firefox 153 and Chrome 151

Rachel Andrew summarizes July's stable releases. Firefox 153 widens how the select element is parsed so that any nested element is allowed inside it, which opens the path to genuinely customizable dropdowns without JavaScript replacements. The same release brings the Picture-in-Picture API to desktop, the Intl.Locale Info API with getCalendars(), getTimeZones(), and getWeekInfo(), and IndexedDB getAllRecords(). Chrome 151 adds the CSS ruby-overhang property, manual slot assignment in declarative Shadow DOM through the shadowrootslotassignment attribute, and the animation property on the AnimationEvent and TransitionEvent interfaces.

web.devverified

Webmedium signal

July Baseline digest: Intl.Locale and rect() cross the line

The edition published on August 10 moves Intl.Locale into newly available — a standard way to parse, modify, and inspect Unicode language and region tags, useful anywhere dates, numbers, and currencies are formatted without an extra library. Array.fromAsync(), which turns asynchronous data streams into an array, and the CSS rect() function, which describes rectangular shapes for clip-path and offset-path far more readably than going through polygon(), both moved to widely available. An updated MDN guide to choosing image formats shipped alongside the digest.

web.devverified

Agentsmedium signal

Claude Code sessions can now message each other

Two open sessions can talk: Claude finds the others with ListAgents and sends a message through SendMessage, either because you asked or on its own when a change in one session affects the work in another. Only text Claude writes crosses over — never conversation history, never files. Since version 2.1.232 you can also address a session directly by typing @ followed by its name. The feature runs on macOS and Linux and needs version 2.1.224 or later, and /list-agents shows what is available. It pays off most when you keep the frontend and backend of the same project in separate sessions.

Anthropicverified

Modelsstrong signal

Gemini 3.7 Flash: half the price and markedly better on code

Google released Gemini 3.7 Flash on August 13, just three weeks after 3.6. The gain on coding benchmarks is large: the DeepSWE v1.1 score rises from 49.0% to 65.3%, and FrontierCode 1.1 from 34.4% to 43.6%. Introductory pricing, in force through the end of 2026, is $0.75 per million input tokens and $3.75 per million output tokens — half what its predecessor launched at. The model is available through the Gemini API, in AI Studio, Android Studio, and Antigravity, and in Spark for AI Pro and Ultra subscribers.

9to5Google / Googleverified

Agentsmedium signal

Auto mode becomes the default permission mode

Since August 14, auto mode is the default permission mode for new sessions on the Pro, Max, and Team plans. Instead of interrupting at every action, a classifier in the background lets safe actions through and stops risky ones. If you set a default mode yourself, it stays in force until you accept a one-time offer to switch, and a mode an organization mandates does not change automatically. One detail matters for your quota: the classifier calls auto mode makes no longer count against usage limits.

Anthropicverified

Industrymedium signal

OpenAI introduces Private Safety Processing

On August 19, OpenAI released a preview of a system that recognizes abuse patterns across several linked sessions without retaining user content. The idea is that only a narrowly defined safety signal is passed on, without exposing the prompts and responses themselves, which keeps the zero data retention guarantee for paid API customers in force. The target is attackers who split a request across several conversations to avoid detection. The approach is the opposite of Anthropic's policy, which keeps data for 30 days with controlled human review; if you are choosing a vendor for a project with sensitive client data, that difference is worth understanding before you sign. Wider rollout and a technical paper were announced for September.

TechCrunch / OpenAIverified

Agentsweak signal

Claude Code 2.1.235 to 2.1.237: Concise style, ANTHROPIC_DEFAULT_MODEL, and spell checking

Three releases in three days change several things about daily work. Version 2.1.237 introduces a built-in Concise output style, in which Claude leads with the result and skips the preamble and the recap, and fixes prompt caching for sessions that go through a proxy gateway or a custom base URL. Version 2.1.236 adds the ANTHROPIC_DEFAULT_MODEL environment variable, which sets the model new sessions start on, and a notify_when_idle option that asks another session to report when it frees up. Version 2.1.235 introduces optional spell checking through aspell, hunspell, or ispell, and reduces memory and processor use for sessions running in the cloud.

Anthropicverified