AI Newsway

Holo4's Best Computer-Use Model Is Open Weights β€” Just Not for Business

H Company's 27B agent posts 61.7% on OSWorld 2.0 at a seventh of Opus 5.5's cost, but only the weaker Mixture-of-Experts sibling carries an Apache 2.0 licence

|5 min read0
AI Summary
H Company released Holo4 on September 28, two open-weight computer-use agents that operate software through GUIs, code, MCP and APIs. The dense 27B model scores 61.7% on OSWorld 2.0 at an estimated $1.22 per task, well below Opus 5.5's 81.8% but at a seventh of the cost. Its weights are CC BY-NC 4.0, so only the weaker Apache 2.0 Mixture-of-Experts version is free for commercial deployment.
A person working alongside an on-screen assistant β€” the kind of desktop interaction Holo4 is trained to perform on its own.
A person working alongside an on-screen assistant β€” the kind of desktop interaction Holo4 is trained to perform on its own.

Paris-based H Company released Holo4 on September 28, a pair of open-weight computer-use models whose smaller 27-billion-parameter version tops every score the lab published β€” and which cannot legally be shipped inside a commercial product.

The catch sits in the model cards rather than the announcement. Holo4-27B on Hugging Face carries a CC BY-NC 4.0 licence, restricting it to non-commercial use. Only the weaker sibling, Holo4-35B-A3B, is Apache 2.0 β€” and on long desktop workflows it scores roughly half as much.

Key takeaways

  • Holo4-27B scores 61.7% on OSWorld 2.0 at an estimated $1.22 per task, against 81.8% at $8.48 for Claude Opus 5.5.
  • The 27B weights are CC BY-NC 4.0; only the 35B-A3B Mixture-of-Experts version is Apache 2.0, and it reaches 30.9% on the same benchmark.
  • H trained both models on about 10,000 tasks generated by its own Agentic Task Factory and published every benchmark trajectory for replay.

One model for GUIs, code and tool calls

Holo4 is built to drive software through whatever interface is in front of it: clicking and typing on a screen, writing and running its own code, or calling MCP and API tools. In its release post, H argues that most agentic models pick one lane β€” GUI-trained models are blind without a screen, while tool-calling models stall in front of an application that exposes no API β€” even though a single business task routinely needs both.

The same checkpoint runs on desktops, the web, Android, a code sandbox and against business APIs, and is called identically in each case. Both sizes went live on the H Models API on launch day, alongside an updated small model called Holotron4 Nano.

Where the numbers actually land

On the original OSWorld desktop suite, Holo4 27B reaches 85.2% at $0.08 per task, narrowly ahead of its own base model Qwen3.8 27B at 84.3% and $0.22, and just behind Fable 5 at 86.0%. The margin widens on OSWorld 2.0, the harder set of long workflows, where Holo4 27B's 61.7% partial score and 41.5% full-success rate beat Qwen3.8 27B's 48.0% and 19.4% at roughly a third of the cost.

It does not beat the frontier. Opus 5.5 leads OSWorld 2.0 at 81.8% and the ALE-CLI expert-workflow split at 63.7% against Holo4's 44.1%, and Opus 5 holds AutomationBench at 50.3% against 45.4%. H's pitch is the cost column beside those scores: $0.05 per AutomationBench task versus $3.05. On AndroidWorld the 27B model posts 85.1%, behind Fable 5's 88.8%.

How it was trained

H's Agentic Task Factory builds interactive environments and verifiable tasks from documentation alone β€” product docs and screenshots of real websites or open-source software. The lab says it has produced about 10,000 tasks so far, split roughly into 4,000 web-app tasks, 3,000 MCP-server tasks and 3,000 desktop and OS tasks.

Supervised fine-tuning ran on 127 billion tokens, about three-quarters of it successful agentic trajectories weighted heavily toward desktop work. Asynchronous online reinforcement learning then trained two specialised LoRA experts β€” one for desktop and web, one for terminal, MCP and API β€” which were merged back into the fine-tuned model at equal weight with no further training.

H also rebuilt the harness around the model, lifting its ceiling from 200 actions over two hours to 500 actions over six. The largest gains, according to the lab, came from giving the agent durable memory across hundreds of steps and a shell on the desktop machine itself.

What to watch next

Every trajectory behind the published scores is downloadable, an unusual disclosure for a lab quoting benchmark wins and one that invites the independent replay that Alibaba's Qwen3.8 open release did not offer. Optimised DSpark drafter checkpoints are promised within days. For teams that need commercial rights today, though, the practical choice is the hosted API or the Apache-licensed model that scores half as well.

FAQ

Can Holo4 be used in a commercial product?

Only partly. Holo4-35B-A3B is released under Apache 2.0 and can be deployed commercially, but the higher-scoring Holo4-27B is CC BY-NC 4.0 and limited to non-commercial use. Both models are also offered through the paid H Models API, which carries no such restriction.

How does Holo4 compare with Claude Opus 5.5 on computer use?

It trails on accuracy and wins on cost. Opus 5.5 scores 81.8% on OSWorld 2.0 against Holo4 27B's 61.7%, but H estimates $8.48 per task for Opus 5.5 versus $1.22 for its own model. H takes the Opus figures from Anthropic's published max-effort runs rather than measuring them in its own harness.

What are Holo4's base models?

The dense 27B model is fine-tuned from Qwen3.8-27B and the Mixture-of-Experts version from Qwen3.6-35B-A3B, according to their Hugging Face model cards. H benchmarks both against those bases, and the gains are largest on long multi-step workflows rather than single-screen tasks.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Qwen-Image-2.1 Puts Transparent Image Editing in 7B Parameters. The License Blocks Commercial Use.
AI & Machine Learning

Qwen-Image-2.1 Puts Transparent Image Editing in 7B Parameters. The License Blocks Commercial Use.

Alibaba's Qwen team released Qwen-Image-2.1 on September 20, an open-weight model that handles text-to-image generation and image editing in a single checkpoint...

Seung Jung8 days ago
Xiaomi's MiMo-V2.6-Pro Is the New Open-Weights Leader, and Its Predecessor Scored 19 on DeepSWE
AI & Machine Learning

Xiaomi's MiMo-V2.6-Pro Is the New Open-Weights Leader, and Its Predecessor Scored 19 on DeepSWE

Xiaomi released MiMo-V2.6-Pro's weights under MIT licence with a 46 on the Artificial Analysis index β€” top among open models. Its predecessor scored 19.0 on DeepSWE; this one scores 71.9.

Seung Jung7 days ago
Told to Fix a Bug, a Coding Agent Retrained and Replaced Its Own Model
AI & Machine Learning

Told to Fix a Bug, a Coding Agent Retrained and Replaced Its Own Model

AI security lab Irregular gave a Qwen3.5-27B agent a maintenance task. It fine-tuned and redeployed the model powering both the app and itself.

Seung Jung9 days ago
Gemini Broke Into Three Outside Systems in May. Google Disclosed It in September.
AI & Machine Learning

Gemini Broke Into Three Outside Systems in May. Google Disclosed It in September.

Google confirmed Gemini accessed three outside systems during a May evaluation, guessing one set of credentials and finding two others in a public repository.

Seung Jung10 days ago
Anthropic Is Running a Physical Biology Lab in the Bay Area
AI & Machine Learning

Anthropic Is Running a Physical Biology Lab in the Bay Area

Anthropic confirmed it runs a Bay Area wet lab for physical biology experiments, and is testing whether Claude can direct robotic systems through them.

Seung Jung10 days ago
A Second Model Broke an Enigma Message. This Time a Human Did More of the Work
AI & Machine Learning

A Second Model Broke an Enigma Message. This Time a Human Did More of the Work

A second frontier model has broken a previously unsolved Enigma message, and the more interesting detail is how much human help it needed. Cybersecurity executi...

Seung Jung3 days ago