Holo4 models add GUI, code, API tools on web, desktop, Android, APIs

H Company's Holo4 models collapse the divide between GUI agents and tool-calling chat models into a single checkpoint.

By Central
Holo4 is a vision-language model that handles desktop, web, Android, code, and API tasks in one framework.
Highlights
  • Holo4 35B-A3B ships under Apache 2.0 license for commercial self-hosting.
  • The per-task cost of Holo4 27B on OSWorld is $0.08, compared to $8-$9 for frontier models.
  • Holo4 models can click buttons, type fields, run code, and call APIs without a separate orchestration model.

Holo4 models add GUI, code, and API tools on web, desktop, Android, and APIs in a single vision-language checkpoint. The new family from H Company collapses the usual divide between screen-driving computer-use agents and tool-calling chat models: the same weights can click a button, type in a field, run code, talk to an MCP server, and call a business API without a separate orchestration model.

H Company has released Holo4, a family of generalist computer-use models designed for AI agents. One set of weights handles the visual world of desktop, web, and mobile interfaces; another skill track handles code, terminal, Model Context Protocol (MCP) servers, and traditional APIs. The release is aimed squarely at the messy reality that modern business software lives in browsers, native apps, legacy desktops, and backend systems all at once.

When a model can drive a desktop app, call a business API, and run code in the same loop for a few cents per task, the conversation shifts from “can an agent do it” to “where do we point it first.”

Holo4 Models Add GUI, Code, and API Tools on Web, Desktop, Android, and APIs

Holo4 ships in two sizes. Holo4 27B is a dense model fine-tuned from Qwen3.8-27B. Holo4 35B-A3B is a Mixture-of-Experts model built on Qwen3.6-35B-A3B and activates only 3B parameters per token. Both support a 256K context window and are served on the H Models API.

The deployment story is practical. Holo4 35B-A3B ships under the Apache 2.0 license, so commercial teams can self-host it. Holo4 27B weights use CC BY-NC 4.0, which means commercial use of the dense model runs through the H Models API. H Company also publishes an open harness, hai-agents, that pairs with both models to handle the recursive agent loop: observe the screen, decide an action, execute the click or tool call, read the result, and repeat.

What Is Holo4?

Holo4 is a vision-language model built specifically for computer use. It is not a chatbot that happens to print code, and it is not a robotic-process-automation script. It sees a screen—whether on a desktop, in a browser, or on Android—and it translates what it sees into actions. Those actions can be clicks, typed text, code execution, MCP tool calls, or API requests.

The architecture matters because the AI agent market has been split into two failure modes. GUI-only agents fall apart when there is no screen to observe. Tool-calling agents stall when an application has no API. Holo4 directly targets that gap: it runs on desktop, web, Android, code sandboxes, and backend APIs, and it can move fluidly between them. A workflow can start with reading a web dashboard, continue by opening a legacy desktop app to retrieve a report, and finish by writing the result into a CRM through a Python script.

Benchmarks: Close to the Frontier, at a Fraction of the Cost

H Company publishes benchmark numbers that put Holo4 in the conversation with frontier models while keeping per-task costs under a dollar in most categories. The results are strongest on OSWorld, where Holo4 27B scores 85.2 percent at roughly $0.08 per task.

OSWorld

Model Score Cost per task
Holo4 27B 85.2 $0.08
Holo4 35B-A3B 80.8 $0.05
Qwen3.8 27B (base) 84.3 $0.22
Fable 5 86.0 n/a
GPT-5.5 78.7 n/a

The benchmark table lists base-model Qwen3.8 27B at 84.3 percent on OSWorld, meaning Holo4 27B improves on its base by less than a point but cuts per-task cost from $0.22 to $0.08. Holo4 35B-A3B achieves 80.8 percent at just $0.05 per task, an aggressive trade-off for teams optimizing on inference spend.

OSWorld 2.0: Long-Horizon Workflows

Holo4 remains competitive on standard OSWorld but loses more ground on OSWorld 2.0, which emphasizes long computer workflows and partial completion. The 27B model finishes 61.7 percent of the work on average, while 35B-A3B falls to 30.9 percent. Frontier models still lead here, but at much higher cost.

Model Score Cost per task
Holo4 27B 61.7 $1.22
Holo4 35B-A3B 30.9 $0.61
Qwen3.8 27B (base) 48.0 $3.49
Claude Opus 5.5 81.8 $8.48
GPT-6 Astra 73.5 $9.07

The long-horizon results reveal the main difference between Holo4 and the frontier. Holo4 27B nearly halves the cost of the base Qwen model while increasing the score from 48.0 to 61.7. Still, the frontier models remain more reliable on tasks that require many steps and continuous self-correction.

AutomationBench: The Holdout Asterisk

On AutomationBench, Holo4 27B posts 45.4 percent at $0.05 per task, while 35B-A3B posts 34.5 percent at $0.02 per task. The efficient economics continue. However, H Company is transparent about an important caveat: 480 of AutomationBench’s 600 public tasks sit in the split the company used for training data collection. On the 120 held-out tasks, Holo4 27B scores 49.3 percent and Holo4 35B-A3B scores 31.7 percent.

Model Score Cost per task
Holo4 27B 45.4 $0.05
Holo4 35B-A3B 34.5 $0.02
Qwen3.8 27B (base) 40.3 $0.09
Claude Opus 5.5 50.3 $3.05
Kimi K3 46.7 $0.43
GPT-5.6 Sol 45.8 $0.67

AndroidWorld and ALE-CLI

On AndroidWorld, a benchmark built around phone apps, Holo4 27B reaches 85.1 percent at $0.08 per task, ahead of GPT-5.6 Sol at 77.6 percent and near Fable 5 and Qwen3.8 Max. On ALE-CLI, an expert Linux command-line benchmark, Holo4 27B scores 44.1 percent, which beats its base but trails frontier models by a wide margin.

Model AndroidWorld Cost per task
Holo4 27B 85.1 $0.08
Holo4 35B-A3B 77.6 $0.07
Qwen3.8 27B (base) 81.9 $0.13
Fable 5 88.8 n/a
Qwen3.8 Max 85.3 n/a
GPT-5.6 Sol 77.6 n/a
Model ALE-CLI Cost per task
Holo4 27B 44.1 $0.82
Holo4 35B-A3B 30.9 $0.29
Qwen3.8 27B (base) 43.5 n/a
Claude Opus 5.5 63.7 $8.22
GPT-6 Astra 61.4 $5.31
Muse Spark 1.3 57.5 $2.44

The costs in these tables are based on H Models API rates for Holo4, provider-published list prices for frontier models, and public API pricing for Qwen base checkpoints. Across every table, the pattern is the same: Holo4 compresses the price of computer-use automation while staying close enough to frontier accuracy to matter for production workloads.

Inside the Holo4 Build

H Company built Holo4 with a training pipeline that treats the agent loop as the core task, not a side effect. The work is organized around four components.

Agentic Task Factory

H Company’s internal pipelines generate environments and verifiable tasks from documentation, screenshots, and real software. The task factory has produced roughly 10,000 tasks: about 4,000 web applications, 3,000 MCP servers, and 3,000 desktop and operating-system environments. This is not a handful of hand-written demos; it is a continuous data engine designed to keep the model honest across many interfaces.

Supervised Fine-Tuning

The SFT set contains 127 billion tokens. About three-quarters of that data is successful agentic trajectories. By interface, the distribution is roughly 45 percent desktop, 14 percent web, 12 percent MCP and API, and 3 percent mobile. The remaining 26 percent covers multimodal reasoning, GUI grounding, and text-only tool use and coding. The mix means Holo4 spends most of its training time on the difficult task of turning visual observations into reliable actions.

Two RL Experts and a Merge

After SFT, H Company runs asynchronous online reinforcement learning to train two LoRA experts. Expert A specializes in desktop and web agents. Expert B handles terminal, MCP, and API workflows. At the end of training, both experts are merged back into the fine-tuned model with equal weight and no further training. The result is one Holo4 checkpoint that can operate across both graphical and tool-based environments without a router model deciding which specialist should run.

Harness Rebuild

H Company also rebuilt its agent harness using failure analysis from OSWorld 2.0. The harness is responsible for the interaction loop: computer vision, action parsing, tool execution, and result tracking. Because the company publishes every trajectory at trajectories.hcompany.ai and through a Hugging Face dataset, teams can study exactly where Holo4 makes mistakes and how it recovers.

Holotron4 Nano

In addition to Holo4, H Company released Holotron4 Nano, a smaller model built on NVIDIA’s Nemotron 3 Nano Omni through the Nemotron Coalition. Holotron4 Nano extends the same computer-use design into a nano-class footprint for edge and low-latency deployments. The existence of a planned Nano variant signals that H Company intends Holo4 to run not just in the cloud, but on local hardware where privacy, cost, or latency rules out hosted APIs.

Pricing, Hosting, and Local Inference

Holo4’s API pricing undercuts most frontier hosted agents. Holo4 27B costs $0.40 per million input tokens and $3.00 per million output tokens. Holo4 35B-A3B costs $0.30 and $2.00, respectively. The API is OpenAI-compatible at https://api.hcompany.ai/v1, which makes it a drop-in replacement for existing agent frameworks.

The open-weight story is just as important. Holo4 35B-A3B is available under Apache 2.0, so enterprises can self-host it for commercial workloads. Holo4 27B is CC BY-NC 4.0, meaning non-commercial self-hosting is allowed while commercial use flows through the H Models API. Both models are available in BF16, FP8, NVFP4, and 4-bit GGUF formats. H Company documents local inference with vLLM and llama.cpp, so teams can run the MoE model with only 3B active parameters on modest hardware.

What Holo4’s Release Signals for the Agent Market

The significance of Holo4 goes beyond benchmark scores. H Company is making a strategic bet that agentic AI does not need a single model that is perfect at everything. Instead, the future is a model that can move between the GUI and the terminal, between a web page and a REST API, without hand-written glue code.

That bet challenges the assumption that the only viable agent architectures are vast frontier deployments with custom tooling for every environment. By releasing a 3B-active-parameter model under Apache 2.0, H Company is targeting developers who want to run a local agent that can see the screen and call tools without sending screenshots to a cloud service. The 35B-A3B model is small enough to deploy on a single workstation or a mid-range GPU node, yet it demonstrates competitive performance on several benchmarks.

For enterprises, the practical consequence is a lower cost of scaling computer-use agents. The per-task cost of Holo4 27B on OSWorld is $0.08. The same task on Claude Opus 5.5 and GPT-6 Astra costs between $8 and $9. Even when those frontier models score higher, the price-performance crossover favors Holo4 for high-volume processes such as data entry, screen scarping, report generation, and business-application automation.

The transparency around automation data is also notable. H Company does not hide the fact that most AutomationBench public tasks are in its training distribution. Publishing every trajectory makes the model’s failures auditable rather than leaving teams to discover blind spots after deployment.

Holo4 models add GUI, code, and API tools on web, desktop, Android, and APIs at a moment when the industry is moving from chatbots to agents that actually complete tasks. The remaining gap on long-horizon workflows shows that the frontier is still ahead, but the price difference is so large that self-hosted teams will find it hard to ignore. When a model can drive a desktop app, call a business API, and run code in the same loop for a few cents per task, the conversation shifts from “can an agent do it” to “where do we point it first.”

Questions answered
  • What is Holo4?Holo4 is a vision-language model built specifically for computer use, capable of seeing screens and translating them into actions.
  • What sizes does Holo4 come in?Holo4 ships in two sizes: Holo4 27B and Holo4 35B-A3B, with the latter activating only 3B parameters per token.
  • What is the cost advantage of Holo4?Holo4 27B costs $0.08 per task on OSWorld, while frontier models cost between $8 and $9 per task.
  • Is Holo4 open source?Holo4 35B-A3B is under Apache 2.0 license, allowing commercial self-hosting, while Holo4 27B uses CC BY-NC 4.0.
  • What is hai-agents?hai-agents is an open harness that pairs with Holo4 models to handle the recursive agent loop of observing, deciding, and executing actions.
Share This Article