Paris-based H Company released Holo4 on September 28, a pair of open-weight computer-use models whose smaller 27-billion-parameter version tops every score the lab published β and which cannot legally be shipped inside a commercial product.
The catch sits in the model cards rather than the announcement. Holo4-27B on Hugging Face carries a CC BY-NC 4.0 licence, restricting it to non-commercial use. Only the weaker sibling, Holo4-35B-A3B, is Apache 2.0 β and on long desktop workflows it scores roughly half as much.
Key takeaways
- Holo4-27B scores 61.7% on OSWorld 2.0 at an estimated $1.22 per task, against 81.8% at $8.48 for Claude Opus 5.5.
- The 27B weights are CC BY-NC 4.0; only the 35B-A3B Mixture-of-Experts version is Apache 2.0, and it reaches 30.9% on the same benchmark.
- H trained both models on about 10,000 tasks generated by its own Agentic Task Factory and published every benchmark trajectory for replay.
One model for GUIs, code and tool calls
Holo4 is built to drive software through whatever interface is in front of it: clicking and typing on a screen, writing and running its own code, or calling MCP and API tools. In its release post, H argues that most agentic models pick one lane β GUI-trained models are blind without a screen, while tool-calling models stall in front of an application that exposes no API β even though a single business task routinely needs both.
The same checkpoint runs on desktops, the web, Android, a code sandbox and against business APIs, and is called identically in each case. Both sizes went live on the H Models API on launch day, alongside an updated small model called Holotron4 Nano.
Where the numbers actually land
On the original OSWorld desktop suite, Holo4 27B reaches 85.2% at $0.08 per task, narrowly ahead of its own base model Qwen3.8 27B at 84.3% and $0.22, and just behind Fable 5 at 86.0%. The margin widens on OSWorld 2.0, the harder set of long workflows, where Holo4 27B's 61.7% partial score and 41.5% full-success rate beat Qwen3.8 27B's 48.0% and 19.4% at roughly a third of the cost.
It does not beat the frontier. Opus 5.5 leads OSWorld 2.0 at 81.8% and the ALE-CLI expert-workflow split at 63.7% against Holo4's 44.1%, and Opus 5 holds AutomationBench at 50.3% against 45.4%. H's pitch is the cost column beside those scores: $0.05 per AutomationBench task versus $3.05. On AndroidWorld the 27B model posts 85.1%, behind Fable 5's 88.8%.
How it was trained
H's Agentic Task Factory builds interactive environments and verifiable tasks from documentation alone β product docs and screenshots of real websites or open-source software. The lab says it has produced about 10,000 tasks so far, split roughly into 4,000 web-app tasks, 3,000 MCP-server tasks and 3,000 desktop and OS tasks.
Supervised fine-tuning ran on 127 billion tokens, about three-quarters of it successful agentic trajectories weighted heavily toward desktop work. Asynchronous online reinforcement learning then trained two specialised LoRA experts β one for desktop and web, one for terminal, MCP and API β which were merged back into the fine-tuned model at equal weight with no further training.
H also rebuilt the harness around the model, lifting its ceiling from 200 actions over two hours to 500 actions over six. The largest gains, according to the lab, came from giving the agent durable memory across hundreds of steps and a shell on the desktop machine itself.
What to watch next
Every trajectory behind the published scores is downloadable, an unusual disclosure for a lab quoting benchmark wins and one that invites the independent replay that Alibaba's Qwen3.8 open release did not offer. Optimised DSpark drafter checkpoints are promised within days. For teams that need commercial rights today, though, the practical choice is the hosted API or the Apache-licensed model that scores half as well.
FAQ
Can Holo4 be used in a commercial product?
Only partly. Holo4-35B-A3B is released under Apache 2.0 and can be deployed commercially, but the higher-scoring Holo4-27B is CC BY-NC 4.0 and limited to non-commercial use. Both models are also offered through the paid H Models API, which carries no such restriction.
How does Holo4 compare with Claude Opus 5.5 on computer use?
It trails on accuracy and wins on cost. Opus 5.5 scores 81.8% on OSWorld 2.0 against Holo4 27B's 61.7%, but H estimates $8.48 per task for Opus 5.5 versus $1.22 for its own model. H takes the Opus figures from Anthropic's published max-effort runs rather than measuring them in its own harness.
What are Holo4's base models?
The dense 27B model is fine-tuned from Qwen3.8-27B and the Mixture-of-Experts version from Qwen3.6-35B-A3B, according to their Hugging Face model cards. H benchmarks both against those bases, and the gains are largest on long multi-step workflows rather than single-screen tasks.






