French dev aims to solve bots' blindness so they can understand GUIs
Because sometimes escaping your sandbox requires clicking a button
Most LLMs are great at answering prompts, but fall short when it comes to navigating around the desktop in Windows or Linux. French AI model dev H unveiled a pair of computer use models on Monday aimed at handling graphical user interfaces (GUIs).
Throughout computing history, computer use largely falls into three categories: command line interfaces (CLIs), application programming interfaces (APIs), and GUIs. AI agents can easily plug into the first two, but navigating desktop environments and applications that often prioritize form before function remains an ongoing challenge.
H's Holo 4 family of models aims to address this challenge by enabling relatively small but capable models to tackle all three computer use scenarios including pointing, clicking, scrolling, and typing their way through graphical interfaces originally meant for us meatbags.
REG AD
Fine tuned using supervised training and reinforcement learning, Holo 4 is built atop Alibaba's Qwen 3.8 27B and Qwen 3.6 35B-A3B models. And by optimizing for CLIs, APIs, and GUIs, H claims that its models achieve far greater versatility than pure computer use models might otherwise. Alongside Holo 4, H has also updated its Holotron model, which is based on Nvidia's Nemotron 3, with similar capabilities.
REG AD
In one example, the company showed Holo 4 27B taking advantage of FreeCAD's macro function to programmatically design a 3D model of the Eiffel Tower rather than manually building it using primitives like cubes. In another demo, H did the opposite using extruded shapes to recreate the company's logo, showing the model's flexibility.
As with any model dev's benchmarks, take these claims with a grain of salt, but if H is to be believed, Holo 4 outperforms significantly larger frontier models from the likes of OpenAI, while using a fraction of the parameters.
Curiously, this doesn't mean that they're cheaper. In fact, while the company shows higher scores, in many cases the models end up costing more per task. Given what we know about Qwen 3.8 27B, this is likely due to Holo 4 using substantially more "thinking" tokens in order to arrive at a final result relative to something like GPT 6 Luna, which doesn't perform as well in the OSWorld 2.0 benchmark, but costs substantially less.
Having said that, the open weights models' diminutive size means that researchers, AI enthusiasts, and enterprises should be able to run them on relatively modest hardware. A 24 GB Nvidia RTX 3090 should be more than capable of running these models at 4-bit precision.
H clearly expects users to do just that since alongside BF16, FP8, and NVFP4 weights, it's also made a Llama.cpp (and by extension LM Studio and Ollama)-friendly GGUF version of the model available for download on Hugging Face.
The model dev says that it also plans to release DSpark draft weights in order to speed up inference using a technique called speculative decoding. We've explored this performance-enhancing inference tech in the past, but in a nutshell it uses a small model to guess the outputs of a larger model. When it works, users experience a speedup in token processing and generation and, when it doesn't, it falls back to the base model ensuring no loss in output quality.
However, the models aren't worth much without a harness. H has developed several agentic harnesses including its open source HAI-Agents harness, which is available for download on its GitHub. However, in theory the models should work with third-party computer use harnesses.
H isn't the only model dev focused on computer use applications. At AWS' Re:Invent conference last year, the company announced its own set of computer-use models. Meanwhile, the big three American model labs, OpenAI, Google, and Anthropic, are also investing in this capability, perhaps because escaping their sandbox sometimes requires pushing a button. ®
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.