Skip to main content

Holo4 Computer-Use Models: What Is Open, What Works, and What to Test

4 min read

H Company released Holo4 models for GUI, code and API work. The 27B checkpoint is downloadable, but its license and benchmark caveats matter.

Holo4 Computer-Use Models: What Is Open, What Works, and What to Test

Computer-use agents often break at the boundary between a screen, a terminal, and an API. H Company’s Holo4 release is designed to cross those boundaries with the same model. Its downloadable weights enable independent testing, but the license and the harness behind the headline score deserve a close read.

H Company introduced Holo4 on September 28, 2026. The series includes a 27B dense model and a 35B-A3B mixture-of-experts model, both offered through the H Models API. The company also published model weights and a smaller Holotron4 Nano release. Holo4 is designed to click and type in graphical interfaces, write and run code, and use MCP or API tools when they are available.

Why the cross-interface approach matters

A real business task rarely stays in one interface. An agent may read a dashboard, export a file, check a database through a tool, and then update a record. A GUI-only model loses efficiency when an API exists; an API-only model cannot handle a program that exposes only a screen. H says Holo4 uses the same model across desktop, web, Android, code sandboxes, and business APIs.

That is a model-design claim, not permission to connect it to every system. The execution harness, identity, network access, and approval policy still determine what an agent can do. Our agent tool-access guide explains the difference between granting a connection and authorizing a specific object-level action. The computer-use approval guide applies the same principle to screen-driven work.

The benchmark claim, with its limits

H reports that Holo4 27B reaches 61.7% on OSWorld 2.0, while its 35B-A3B model reaches 30.9%. It compares the 27B result with higher scores from much larger closed models and describes a lower estimated cost per task. H also publishes replayable trajectories. This transparency is useful, but it does not make every comparison apples-to-apples.

H’s methodology notes that releases, harnesses, and task subsets differ across plotted systems. Some cost estimates come from H’s API rates and token counts, while reference points use other providers’ prices or public leaderboards. A deployment decision should therefore compare agents under the same tasks, environment, action limits, and success rubric. Record failures and retries, not only final benchmark percentages.

Are Holo4’s weights open source?

H makes the weights available for download in several formats, including BF16, FP8,8, and GGUF. However, the Holo4-27B model card lists a Creative Commons Attribution-NonCommercial 4.0 license. That distinction matters: open weights are not the same as unrestricted permission to use the model in a paid product. Organizations considering commercial self-hosting should review the exact license for each checkpoint and ask H for the appropriate terms. API access has its own service terms and pricing.

The published model card also lets developers examine the checkpoint and run their own local tests. But a local model is only one component of a computer-use agent. H says it rebuilt the execution harness to retain memory across long runs and to provide a shell on the desktop machine. Those choices can materially affect results; document them alongside the model version in any comparison.

A practical evaluation before deployment

  1. Choose a reversible workflow. Start with a sandbox account and a task whose correct final state you can check automatically.
  2. Fix the environment. Keep the same app version, screen size, available tools, and data for each candidate agent.
  3. Record the path. Save screenshots, tool calls, tokens, elapsed time, retries, and human interventions.
  4. Test boundaries. Include an action that must be refused or sent for approval, such as changing account permissions or sending external data.
  5. Review the license and operating cost. Separate API use from self-hosting rights, hardware requirements, and support terms.

Holo4’s meaningful contribution is a downloadable agent model intended to move between several software interfaces, backed by inspectable trajectories. Its usefulness to a team will depend on the workflow, the harnesses,s and the permission boundary. A strong public score is a starting point for that test, not the result.

Primary sources

Leave a comment

Your email address will not be published. Required fields are marked *