ESPN DeportesNBA 2026-27: Giannis, LaMelo, LeBron: Los 8 cambios de equipo más interesantesInquirerHighlights: Day 30 of Sara Duterte impeachment trial | Sept. 28, 2026Daily MaverickEAST RAND MURDERS: After another woman murdered in Ekurhuleni, here’s what we know about the killingsThe Jerusalem PostTrump vows victory in Iran war 'very soon,' officials in talks with Iranian mediatorsESPNJ.J. McCarthy isn't the answer for the Giants ... right? Barnwell on the trade's potential impactSouth China Morning PostIn the US-China AI race, the real fight is keeping humans in controlBillboardElla Langley’s ‘Choosin’ Texas’ Extends Hot 100 Record With 24th Week at No. 1HipertextualGemini cambia para siempre: su función más útil dejará de existir en noviembreCBS SportsNFL asks Homeland Security to remove Brian Dawkins highlight video from social mediaTechCrunchSource: Inference provider Modal Labs closing in on $750M round at $15.75B valuationسكاي نيوز عربيةأوليسيه يمنح فرنسا فوزا قاتلا على بلجيكا في دوري الأممPremium TimesPolice arrest Delta commissioner, four others over deputy director’s death
The Daily Newsstand · Free, Always
Monday, September 28, 2026

French dev aims to solve bots' blindness so they can understand GUIs

Translate

Because sometimes escaping your sandbox requires clicking a button

Most LLMs are great at answering prompts, but fall short when it comes to navigating around the desktop in Windows or Linux. French AI model dev H unveiled a pair of computer use models on Monday aimed at handling graphical user interfaces (GUIs).

Throughout computing history, computer use largely falls into three categories: command line interfaces (CLIs), application programming interfaces (APIs), and GUIs. AI agents can easily plug into the first two, but navigating desktop environments and applications that often prioritize form before function remains an ongoing challenge.

H's Holo 4 family of models aims to address this challenge by enabling relatively small but capable models to tackle all three computer use scenarios including pointing, clicking, scrolling, and typing their way through graphical interfaces originally meant for us meatbags.

REG AD

Fine tuned using supervised training and reinforcement learning, Holo 4 is built atop Alibaba's Qwen 3.8 27B and Qwen 3.6 35B-A3B models. And by optimizing for CLIs, APIs, and GUIs, H claims that its models achieve far greater versatility than pure computer use models might otherwise. Alongside Holo 4, H has also updated its Holotron model, which is based on Nvidia's Nemotron 3, with similar capabilities.

REG AD

In one example, the company showed Holo 4 27B taking advantage of FreeCAD's macro function to programmatically design a 3D model of the Eiffel Tower rather than manually building it using primitives like cubes. In another demo, H did the opposite using extruded shapes to recreate the company's logo, showing the model's flexibility.

As with any model dev's benchmarks, take these claims with a grain of salt, but if H is to be believed, Holo 4 outperforms significantly larger frontier models from the likes of OpenAI, while using a fraction of the parameters.

Curiously, this doesn't mean that they're cheaper. In fact, while the company shows higher scores, in many cases the models end up costing more per task. Given what we know about Qwen 3.8 27B, this is likely due to Holo 4 using substantially more "thinking" tokens in order to arrive at a final result relative to something like GPT 6 Luna, which doesn't perform as well in the OSWorld 2.0 benchmark, but costs substantially less.

Having said that, the open weights models' diminutive size means that researchers, AI enthusiasts, and enterprises should be able to run them on relatively modest hardware. A 24 GB Nvidia RTX 3090 should be more than capable of running these models at 4-bit precision.

H clearly expects users to do just that since alongside BF16, FP8, and NVFP4 weights, it's also made a Llama.cpp (and by extension LM Studio and Ollama)-friendly GGUF version of the model available for download on Hugging Face.

The model dev says that it also plans to release DSpark draft weights in order to speed up inference using a technique called speculative decoding. We've explored this performance-enhancing inference tech in the past, but in a nutshell it uses a small model to guess the outputs of a larger model. When it works, users experience a speedup in token processing and generation and, when it doesn't, it falls back to the base model ensuring no loss in output quality.

However, the models aren't worth much without a harness. H has developed several agentic harnesses including its open source HAI-Agents harness, which is available for download on its GitHub. However, in theory the models should work with third-party computer use harnesses.

H isn't the only model dev focused on computer use applications. At AWS' Re:Invent conference last year, the company announced its own set of computer-use models. Meanwhile, the big three American model labs, OpenAI, Google, and Anthropic, are also investing in this capability, perhaps because escaping their sandbox sometimes requires pushing a button. ®

View the original on The Register →

KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.