Skip to content
Modelsmedium signalverified

H Company releases Holo4, open-weight models for agents that operate a computer

H Company released two Holo4 models on September 28, 2026, with weights on Hugging Face. The larger one scores higher on the company's own tests but may not be used commercially. The smaller one is licensed under Apache 2.0.

By Redakcija WebAiRadarPublished 2 min readwritten by a model
Image: Hugging Face

Source

Holo4: powering generalist computer-use agents

Hugging Face Blog · Original published September 28, 2026

H Company, a French AI lab, released Holo4 on September 28, 2026: a pair of models built to drive software on their own. Holo4 can click and type on a screen, write and run its own code, and call MCP or API tools, and the company says it picks whichever fits the task. The weights for both sizes are on Hugging Face, and both are also served through the company's H Models API.

What was released

The two models differ in architecture as well as size. Holo4-27B is a dense model with 27 billion parameters, built on Qwen3.8. Holo4-35B-A3B is a mixture-of-experts model with 35 billion parameters in total, built on Qwen3.5 MoE. Both configurations allow a context of up to 262,144 tokens.

H Company says the same model runs on desktops, on the web, on Android, in a code sandbox, and against business APIs, and that it is called the same way everywhere. Next to Holo4, the company released Holotron4 Nano, a smaller agent model built on Nvidia's Nemotron 3 Nano Omni.

  • Holo4 weights come in BF16, FP8, NVFP4, and 4-bit GGUF, so the quantized files can go straight into local runtimes that read GGUF.
  • Holotron4 Nano ships only in BF16 and FP8, under the NVIDIA Open Model Agreement.
  • Every agent trajectory behind the published benchmark scores is available as a dataset on Hugging Face.
  • The model cards point to the hai-agents harness, which sends screenshots and tool results to the model and carries out the actions it requests.

What the vendor's numbers show

According to H Company's measurement, Holo4-27B scores 61.7% on OSWorld 2.0, a benchmark for controlling a desktop. Holo4-35B-A3B reaches 30.9%. In the same chart the company lists Opus 5.5 at 81.8%. On AutomationBench, which tests work through APIs, the company reports 45.4% for the larger model and 34.5% for the smaller one.

The company estimates that an OSWorld 2.0 task costs $1.22 with Holo4-27B and $0.61 with Holo4-35B-A3B. Those costs are calculated from the tokens of each run at H Models API rates. H Company itself notes that benchmark releases, task subsets, and harnesses differ between the models it compares, and no independent measurement has been published yet.

The license decides which model you can ship

The two sizes are not licensed the same way. Holo4-27B is released under CC BY-NC 4.0, which does not allow commercial use. Holo4-35B-A3B is released under Apache 2.0, which does.

This suggests a trade-off for anyone building a product. The model you can put into a paid service without a separate agreement is the one that, by the company's own numbers, scores roughly half as much on OSWorld 2.0. The announcement does not list a per-token price for the H Models API.

„Most agentic models are trained for one interface only.“
H Company

Sources

Related

Anthropic logo
Modelsstrong signal

Anthropic releases Claude Sonnet 5.5, with input and output tokens at half the Opus 5.5 price

Claude Sonnet 5.5 keeps Sonnet 5's prices: $2 per million input tokens and $10 per million output tokens. That is half of what Opus 5.5 charges, while Anthropic's own benchmarks put the two models within a few points of each other. Code that turns thinking off needs a change before the switch, because the old setting now returns an error.

Anthropicverified

Better prompt caching for GPT-6
Modelsstrong signal

OpenAI releases GPT-6 Sol and Luna and halves the API price of both tiers

GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, and GPT-6 Luna costs $0.10 and $0.50. OpenAI measures the 50% cut against the promotional prices it charged for GPT-5.6. The launch also brings a reworked prompt cache with discounts of up to 90% on cached input. Every benchmark comparison with Claude comes from OpenAI and uses Opus 5, not Opus 5.5.

OpenAIverified

Anthropic logo
Modelsstrong signal

Anthropic ships Claude Opus 5.5 and cuts token prices by a fifth

Opus 5.5 charges $4 per million input tokens and $20 per million output tokens, 20% below Opus 5. Cache reads, which dominate the bill for agentic and coding work, fall 60% to $0.20 per million. Anthropic says the model works at the level of Claude Fable 5.1 on most tasks while costing 40% less to run than Opus 5. It also arrives with safeguards that reroute most cybersecurity work to Opus 4.8.

Anthropicverified