Holo4: Powering Generalist Computer-use Agents
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

H Company has released Holo4, a series of models designed to complete tasks across graphical interfaces, code, MCP and APIs. The company reports a 61.7% score for Holo4 27B on OSWorld 2.0 and says its benchmark runs cost less than those of larger frontier models, though evaluation methods and task sets vary.

H Company has released Holo4, a new series of computer-use agent models that can interact with software through graphical interfaces, code, MCP and APIs. The company says the models are intended to handle business workflows that may require switching between those methods, and are available in 27B dense and 35B-A3B Mixture of Experts versions through its H Models API.

Holo4 is designed to click and type on screens, write and run code, and call MCP or API tools, selecting an interface based on the task, according to H Company. The company says the same model can run on desktops, the web, Android, code sandboxes and business APIs, without requiring a separate model for each platform. The release also includes Holotron4 Nano, an updated version of Holotron 3.

On the desktop-control benchmark OSWorld 2.0, H Company reports a score of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. The report compares the 27B result with 81.8% for Opus 5.5. H Company says its models use far fewer parameters and cost less per task than leading closed models, but the comparison draws on different releases, harnesses and task subsets, so the scores are not a uniform head-to-head evaluation.

The company says Holo4 was trained through supervised and reinforcement learning on environments and tasks, including material generated by its Agentic Task Factory. It has published benchmark trajectories for replay or download and offers model files in FP16, FP8 and GGUF formats. H Company also presents examples of Holo4 27B completing FreeCAD modeling and Godot game-building prompts, alongside comparisons with the Qwen base model under the same prompt and harness.

At a glance
announcementWhen: Announced in the H Company report; the…
The developmentH Company announced Holo4, a pair of agentic models built to combine screen interaction, code execution and tool calls in business workflows.

One Agent Across Software Interfaces

Many workplace tasks span more than one kind of software control. A person or agent may need to inspect a screen, edit a file in code, then send information through a business API. H Company’s central claim is that one model can handle those steps together, reducing the need to route work among separate GUI and tool-use agents.

If the approach performs reliably outside benchmark settings, it could make automation easier to deploy across applications that expose different interfaces. The release’s reported benchmark results and example trajectories offer some evidence of capability, but they do not establish how often the model completes varied business tasks correctly or how much human supervision is required. Those practical questions matter alongside the company’s cost claims.

From Single-Interface Agents to Holo4

H Company describes Holo4 as building on its previous model series. Its report argues that agents trained for only screen interaction can be limited when an application has no usable interface for that approach, while agents centered on tool calls may be constrained when software has no API. Holo4 is presented as a response to that divide, combining those modes within a single system.

The report places particular emphasis on long workflows and professional software. OSWorld 2.0 tests desktop interaction, while AutomationBench evaluates API use. For AutomationBench v1.0.6, H Company says it measured Holo4 and two Qwen models in its internal harness; it plans to report Holo4 on the benchmark’s private set after evaluation. Comparisons with other models draw on public-set scores and leaderboard cost figures, which use different evaluation sets.

“Real work is not siloed that way, and a single business task can require combining these different approaches.”

— H Company, in its Holo4 report

Benchmark Comparisons Have Limits

The report does not specify a publication date, and its cross-model comparisons are not based on one shared evaluation setup. H Company’s cost estimates use token consumption and listed prices for some models, while other results and cost figures come from separate sources and evaluation runs. The company says it has not yet evaluated Holo4 on AutomationBench’s private set.

The source does not provide independent verification of the reported results, detailed measures of error rates on business tasks, or evidence about performance across customers’ software environments. It also does not quantify the supervision needed, deployment costs beyond the stated task estimates, or how the model handles failures in extended workflows.

Private-Set Evaluation Still Ahead

H Company says it will report Holo4’s results on AutomationBench’s private set after evaluation. Readers can inspect the company’s published benchmark trajectories, model collection and API quickstart as further evidence of how the system behaves. Independent evaluations using consistent tasks and cost accounting would help clarify how its reported performance compares with other agents in practical deployments.

Key Questions

What is Holo4?

Holo4 is H Company’s new series of agentic models for computer-use tasks. The company says it can interact through GUIs, code, MCP and APIs.

Which Holo4 models are available?

The release includes a 27B dense model and a 35B-A3B Mixture of Experts model, both available through the H Models API. H Company also released Holotron4 Nano.

How did Holo4 score on OSWorld 2.0?

H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. It reports 81.8% for Opus 5.5, while noting that comparisons involve differing releases, harnesses and task subsets.

Has Holo4 been evaluated on AutomationBench’s private set?

Not according to the report. H Company says it will publish Holo4 results on the private set after evaluation.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Accelerating GPT-5.6 Sol Ultrafast

OpenAI reportedly advances GPT-5.6 Sol Ultrafast, aiming for faster deployment; details remain undisclosed, raising industry interest.

AI Is Removing The Middle Class Of Software Engineering?

Experts warn AI automation could displace mid-level software engineers, raising concerns about job security and industry shifts.

A Look At Claude Fable 5.1’S AI Index Leadership And The Cost Line Details

Artificial Analysis ranks Claude Fable 5.1 highest on its AI Index, with a score of 66, but notes it costs about 20% more per task due to verbosity.

SenseTime SenseNova U1.5: Pioneering AI Advancements With Open Source Code

SenseTime announces SenseNova U1.5, an 8-billion-parameter unified vision-language model, and releases its training code publicly, marking a strategic shift in AI transparency.