Your recordings are the training set
Because the models train from video, we can train directly on the explainer recordings your team already has, and deliver working agents for those workflows. No annotation project first.
HuzzleWorld is Huzzle Labs’ family of computer-use models, released at 1B, 2B, 4B and 8B parameters and built on our proprietary architecture, our own RL environments, and training from video. On ScreenSpot-Pro, they are state of the art for their size. HuzzleWorld-2B outscores grounding models up to 36 times its size.
HuzzleWorld-2B can take screenshots, return coordinates, click, scroll and type.
The same properties that make HuzzleWorld small make it practical to deploy inside a company rather than behind someone else’s API.
Because the models train from video, we can train directly on the explainer recordings your team already has, and deliver working agents for those workflows. No annotation project first.
A small model needs far less data, compute and time to train than a large one, while matching its accuracy on the workflows you care about. Shorter path from recording to running agent.
1B to 8B fits on hardware enterprises already run. Deploy HuzzleWorld inside your own environment, so screen recordings and workflow data never leave it.
HuzzleWorld (also written Huzzle World) is Huzzle Labs’ family of computer-use models, released at 1B, 2B, 4B and 8B parameters. Each takes a screenshot and an instruction and returns the exact screen coordinate to act on, which is the grounding step underneath any computer-use agent.
60.8% single-pass and 64.2% with an optional two-pass zoom-in. It reaches 74.8% on office software and 70.3% on text-labelled targets. Both settings are the highest recorded for a 2B model on the official leaderboard.
Single-pass, HuzzleWorld-2B outscores grounding models up to 36 times its size, including UI-TARS-72B (38.1%), UGround-v1-72B (34.5%), Qwen2.5-VL-72B-Instruct (53.3%), and H Company’s Holo2-8B (58.9%) and Holo1.5-7B (57.9%).
A GUI grounding benchmark built from full-resolution screenshots, up to 3840 × 1080, of 26 real professional applications. A prediction only counts if it lands inside the ground-truth box. There is no partial credit.
Grounding runs on every step of an agent trajectory, so its cost and latency multiply across a workflow. A 2B model that matches or beats a 72B model changes where computer-use agents can be deployed and what they cost to run.
On our proprietary architecture, in our own long-horizon RL environments, and from video. Learning from screen recordings instead of large hand-annotated datasets means the models can be trained directly on a company’s existing explainer recordings.
Yes. At 1B to 8B the models fit on hardware enterprises already run, so HuzzleWorld can be deployed inside your own environment, with no data leaving the organization.
A benchmark that runs an agent on a real Ubuntu desktop across applications like Chrome, GIMP, LibreOffice, VLC and VS Code, then runs a script to check whether the task was actually completed. Unlike ScreenSpot-Pro it scores end-to-end task completion, not a single click.
It completes 58.6% of tasks, the highest of any model on the leaderboard published at 8B parameters or smaller, and ahead of GUI-Owl-1.5 32B, DeepMiner-Mano-72B and opencua-72b-preview. Across the leaderboard regardless of size it stands 11th of 41 models.
The step that turns an instruction and a screenshot into a screen coordinate: deciding exactly where to click. It is the part of a computer-use agent that connects a plan to an action, and it runs on every step of a workflow.
Deploying computer-use agents, or training your own?
Let’s talk.
Leaderboard data as of 30 July 2026. Ranks are positions among the 72 ScreenSpot-Pro leaderboard entries with a published parameter count and a score of 10 or above. Charts are plotted from 20% upwards; four entries scoring below 20% fall outside the plotted range. Scores are grounding accuracy across the benchmark’s 23 applications and three operating systems; the benchmark measures localisation on a static screen, not end-to-end task completion. Single-pass and zoom-in are recorded as separate leaderboard entries. Figures for HuzzleWorld-1B and HuzzleWorld-4B are our own measurements on the same benchmark, with leaderboard submissions in review. OSWorld Verified figures refresh automatically from the public OSWorld leaderboard and show the best recorded run per model; HuzzleWorld-8B’s own score is Huzzle-reported. OSWorld publishes no parameter count, so the size-class board lists only models whose published name states a size of 8B or smaller, and models of undisclosed size are left out of it.