# Fig.inc > Fig is a frontier AI lab pioneering systems capable of continual learning and autonomous action across complex, open-ended environments. Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### Privacy Policy URL: https://www.fig.inc/privacy/ Last updated: 2026-06-16T08:36:12.000Z **Last updated: June 15, 2026** ## 1\. Introduction **Metarch Corporation** ("Fig," "we," "us," or "our") is an artificial intelligence research company. This Privacy Policy (the "Policy") describes how we collect, use, and disclose your information when you access or use our website and blog at fig.inc, or otherwise interact with us (collectively, the "Services"). For purposes of this Policy, Fig is the controller of your information. By using our Services, you agree to the collection, use, and disclosure of your information as described here. If you do not agree with this Policy, please do not use our Services. If you have questions or wish to exercise your privacy rights, contact us at **contact@fig.inc**. ## 2\. Changes to this Policy We may update this Policy from time to time. If we do, we will post the updated version here with a new "Last updated" date. If we make material changes to how we use or disclose information, we will make reasonable efforts to notify you as required by applicable law. ## 3\. Personal information we collect and how we use it When you use our Services, we collect personal information from the sources described below. In addition to the specific purposes noted, we may use this information in our legitimate interests to provide and improve the Services, communicate with you, maintain the security and integrity of the Services, comply with legal obligations, prevent fraud or misuse, and protect our rights and those of our users and others. **What you provide us** - **Communications information:** your name, email address, and anything else you choose to include when you contact us (for example by email or social media) or when you submit one of our forms — such as investor, partnership, or recruiting inquiries — so that we can respond to and evaluate your inquiry. - **Subscription information:** your email address if you subscribe to updates, so we can send the updates you requested. You can unsubscribe at any time. **What we receive automatically** We and our service providers collect the following when you use our Services, including through cookies and similar technologies. You can configure your browser to disable cookies, though this may affect how the Services work. We collect and use this information in our legitimate interests to operate the Services, run analytics, and understand how the Services are used. - **Device information:** such as device type, identifiers, and IP address. - **Location information:** the general area from which your device accesses the Services, based on information such as your IP address. - **Log data:** information your browser sends automatically, such as browser type and settings. - **Usage data:** We collect information about your use of our Services, such as the content you access on our website, and your activity on our website. Your browser may let you send a "Do Not Track" or similar signal. Our Services are not currently designed to respond to these signals. California residents can exercise the choices described in Section 10. **What we receive from third parties** We may receive personal information from service providers that help us operate the Services (such as hosting, security, and analytics providers), and we may collect information that is publicly available online. ## 4\. How we disclose personal information We may disclose personal information to: - **Vendors and service providers** who help us operate the Services, including for hosting, storage, security, form submissions, and web analytics. - **Affiliates** within our corporate group, to operate our business and provide the Services. - **Professional advisors** such as auditors and law firms, to comply with legal obligations and to protect and defend our rights. - **Parties to a business transaction**, such as in connection with a merger, financing, acquisition, or sale of assets. - **Government authorities or other parties** where necessary to comply with law, respond to legal or regulatory requests, enforce our terms, or protect the safety and security of our business, users, and others. ## 5\. How long we store personal information We keep personal information for as long as reasonably necessary for the purposes described in this Policy. How long depends on factors such as whether we need the information to provide the Services, resolve disputes, comply with legal obligations, or protect our rights, as well as the nature of the information and the risk of harm from unauthorized use or disclosure. ## 6\. How we protect personal information We make reasonable efforts to protect your personal information from loss, misuse, and unauthorized access or disclosure, but we cannot guarantee perfect security. Please take care when sharing information with us. ## 7\. Your personal information rights Depending on where you live, you may have the right to: - Access or know the personal information we hold about you and how it is processed; - Delete your personal information; - Correct inaccurate personal information; - Port a copy of your personal information to a third party; - Restrict or object to certain processing; and - Withdraw your consent, where we rely on consent. These rights are not absolute, and we may decline a request where permitted by law. To make a request, email **contact@fig.inc**. We will not discriminate against you for exercising your rights, and we may need to verify your identity before responding. ## 8\. Children's privacy Our Services are not directed to children under 13, and we do not knowingly collect their personal information. If you believe we have, email **contact@fig.inc** and we will make reasonable efforts to delete it. ## 9\. Third-party links and integrations Our Services may link to or integrate with third parties, such as social media platforms and form providers. We are not responsible for the privacy practices of those third parties, and we encourage you to review their policies and contact them with any questions. ## 10\. Additional U.S. state disclosures Some U.S. states, such as California, require additional disclosures. For this section, "personal information" includes "sensitive personal information" as defined under the California Consumer Privacy Act ("CCPA"). Over the past 12 months we have collected the following **categories** of personal information: - **Identifiers**, such as name, email address, IP address, and device information; - **Internet or network activity**, such as information about your use of the Services; - **Geolocation data**, such as the general location derived from your IP address; and - **Professional or employment-related information**, where you provide it (for example, in a recruiting inquiry). We use this information for the purposes described above (Section 3) and disclose it to the categories of recipients described in Section 4 (service providers and, where required, government authorities). We retain it as described in Section 5. **We do not "sell" or "share" personal information** as those terms are defined under the CCPA, and we have not done so in the prior 12 months. We do not knowingly sell or share the personal information of individuals under 16. ## 11\. Data transfers We are based in the United States. If you use the Services from outside the United States, your information may be transferred to, stored, and processed in the United States or other countries, which may have data-protection rules different from those of your country of residence. ## 12\. How to contact us Please email us at contact@fig.inc if you have any questions or would like to exercise any of your rights as described herein. ## Posts ### Goal Progress Decays with Task Horizon for Frontier VLMs in Interactive 2D Environments URL: https://www.fig.inc/blog/multinet-v2-preview/ Last updated: 2026-08-26T14:10:04.000Z **[Pranav Guruprasad](https://pranavguru.github.io/?ref=fig.inc)1, 2**, **[Sean Rivera](https://www.linkedin.com/in/sean-au-rivera/?ref=fig.inc)2**, **[Helen Lu](https://www.linkedin.com/in/helen-lu-aa60a4b9/?ref=fig.inc)2, 4**, **[Arushi Jain](https://www.linkedin.com/in/jainarushi08/?ref=fig.inc)2**, **[Hangliang Ren](https://www.linkedin.com/in/hangliang-ren-b74bb2194/?ref=fig.inc)2**, **[Harshvardhan Sikka](https://www.harshsikka.com/?ref=fig.inc)1, 2, 3** 1[Fig](https://fig.inc/?ref=fig.inc); 2[Manifold Research Group](https://www.manifoldrg.com/?ref=fig.inc); 3[Georgia Institute of Technology](https://www.gatech.edu/?ref=fig.inc); 4[Tufts University](https://www.tufts.edu/?ref=fig.inc). ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/r1_failure_reels.gif) Relevant links · [Code](https://github.com/ManifoldRG/MultiNet-v2.0?ref=fig.inc) · [Website](https://multinet.ai/?ref=fig.inc) · [Evaluate your model](https://app.notion.com/p/3bf4b1d3c487800596bbe4a150962cc0?ref=fig.inc) · [Cite this](#cite) TL;DR - **We evaluated three highly performant vision-language models on 2D mazes:** 50 Minigrid mazes, with varying path lengths and mechanisms such as key-doors, switch-gates, and decoys that need to be reasoned through in order to reach the final goal. - **Model performance is near zero:** Across 3 major Vision-Language Models - Claude Opus 4.8, Kimi K2.6, and Qwen 3.6-27B - we saw only 6 maze solves across 150 mazes (50 for each model). Additionally, 45 of the 50 mazes were never solved by any model. - **Ablation-based maze design:** We ran ablation experiments across 540 episodes on 15 held-out mazes to finalize the evaluation protocol with quantified justification behind every decision. - **We observed a variety of distinct failure modes:** The models cannot work out mechanisms they have no prior knowledge of, and each fails differently - Claude jams into walls, Kimi retreads seen ground, Qwen turns without progressing. The models that reason the most solve the fewest mazes. Key Sections · [The evaluation environment](#evaluation environment) · [Evaluation setup](#evaluation setup) · [Results and analysis](#results) · [Discussion](#discussion) · [Limitations](#scope) We evaluated 3 highly capable Vision-Language Models (VLMs) on 50 2D mazes. Across the 3 models, only 6 total mazes were solved, and 45 of the 50 mazes were not solved by any model. The size of the mazes ranges from 8x8 to 14x14 grids. An agent has to navigate through walls, avoid decoys and distractors, and operate objects that are to be operated in a specific order (a key that opens a door, a switch that toggles open a gate), in order to reach a goal tile and solve the maze. The action space consists of six actions: turn left, turn right, move forward, pickup, toggle, and done. Apart from the action space and the instruction to complete the maze by reaching the goal tile, the models are not given any other information about the environment. They have to figure out what to do by exploring the environment, acting in it, and observing the changes in the environment. When a benchmark like the one we present records a zero for a model failing an episode, there are multiple possible reasons for the failure: - The model never worked out what an object in the environment does. - It worked out what the objects are, but could not operate them in the right order. - It executed the wrong actions from the action space. - It made mistakes and could not recover from them to get back on the right path. All of these are completely different failure modes, and each of them requires a very different fix. Going down the wrong path in a scenario such as this is expensive in both time and resources spent, and the cost of not identifying the right failure mode is high: a survey of agentic benchmark construction found flaws that mis-estimate performance by up to 100% \[1\]. Separating the modes of failure means changing the task in different ways: running a maze without the switch-gate mechanism, running it with the optimal path length controlled, and so on. Whichever change moves the score helps narrow down on the failure mode of the model or agent. In this report, we describe an environment built to evaluate long-horizon action and causal reasoning capabilities in models, which by the nature of its design also helps separate the different possible failure modes in models. ## Designing an environment that can attribute failure When a model has to act over a long horizon in an environment whose underlying rules it has not been told, where, how, and why does it break? Through our evaluation in this work, we find that models: - **Solve 6 out of 150 episodes.** 45 of the 50 mazes were not solved by any model, and outside the six solves no episode ended within three executable actions of the goal. - **Fail on mazes with a longer path length to the solution.** Mean action progress falls from 0.287 on mazes solvable in 30 moves or fewer to 0.070 on mazes needing 61 or more. - **Are unable to learn how a mechanism works without prior knowledge.** Across the 105 episodes containing a switch and a gate there were no solves at all, while key-door mechanisms account for four of the six solves. - **Fail to solve a majority of the mazes despite reasoning with an extended thinking token budget.** The two models that spent more than 20x the output tokens solved one episode each, against four for the model that spent the fewest. In order to identify and separate these different failure modes, we believe that four properties are needed in an environment: 1. **Knowledge-free:** No pre-required domain knowledge. A failure cannot be explained away as an unfamiliar interface and a human baseline requires no training in order to be an appropriate comparison. 2. **Generative:** If the environment is built as a composition of multiple parameters, rather than entirely designed by hand, there is no static answer key that could get contaminated by frontier model pre-training, and the difficulty scales with the models instead of saturating 3. **Projectable:** Because the environment structure carries no domain content, the underlying task can be rendered in language or in 3D or any other modality. Measuring performance on the same underlying task across modalities is a good measurement of capability instead of familiarity with an interface. 4. **Verifiable:** A method to obtain the exact optimal action sequence from any reachable state yields an objective difficulty scale and the ability to provide partial credit to a trajectory \[2\]. Mazes with causal mechanisms are a simple representative structure carrying the 4 properties discussed above. In MultiNet v2.0-Gridworld, we run evaluations on this substrate in a very basic rendering. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/projection_figure.png) Figure 1: Maze as an underlying substrate allows projection of the same environment and task into multiple domains such as language and 3D simulation ## The Evaluation Environment The underlying task we present in this evaluation is a model in a 2D Minigrid \[3\] maze with an action space containing 6 valid actions: turn left, turn right, move forward, pickup, toggle, and done. A model is required to navigate corridors bounded by walls, decoys and distractors, and mechanisms that need to be understood and operated in a specific order, with an end goal of reaching a target tile. Nothing about the environment is explicitly explained, and needs to be figured out by the model exploring the environment. We chose to keep the environment and its components unexplained because prior domain knowledge is a confounding variable. Measuring capability requires controlling for priors \[4\], and with this design, a failure cannot be attributed to an unfamiliar API, library or interface, because there are none. At each turn, the model is provided with persistent information on where it started in the maze, a prose activity summary of every mechanism event so far, and up to four rendered observation frames: the 3 most recent steps with inventory and the action taken at each, then the current frame with a request for the next action. The model receives no progress signal and no explicit verbal information that a move it took failed. In order to succeed, the model needs to navigate to the goal tile - this is considered a solve. Progress is another metric we track, and is defined as: 1 − dremaining / dstart where dremaining is the remaining oracle distance from where the model ends the episode, and dstart is the total oracle distance from the start state as calculated by a BFS implementation. This is computed over two ways: over tiles, which ignores the barriers, and over executable actions which accounts for them. The action variant makes more sense to rely on because the tile variant credits the model for being close to the goal even if it is sealed behind a door it never opened. We implement 2 termination conditions other than the solve - a step cap, and a stall watchdog. The step cap is an upper limit on the number of steps the model is allowed to take in a given episode, and is defined as 3x the optimal BFS step count. The stall watchdog is implemented to prevent models wandering aimlessly and terminates a model’s episode if there is no novelty in position, inventory, and mechanism state together for 30 steps. This makes the step budget for a given episode relative to the difficulty of the maze instead of a flat action step cap. We observe in our evaluation results that the watchdog terminates almost every episode. Only 11 out of 150 episodes were terminated due to a success or the model hitting the step cap. ### What does each mechanism isolate? ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/maze_figure.png) Figure 2: The maze environments we evaluate models in contain various combinations of mechanisms that are typically required to be solved in a specific order Solving each of the mazes in our evaluation set requires 4 things in sequence - work out how the mechanisms in the environment work, reason about the order in which the mechanisms need to be operated, execute the right actions, and recover in case something goes wrong. Each mechanism is included in the environment because it tests one of the 4, and because each can be added or removed independent of the others - failure modes can be classified accurately. - Keys and doors test whether a model can understand and operate a mechanism it has prior knowledge about. All frontier models have seen keys being used to open doors in their training data and do not have to explore the environment to understand how they work. - Switches and gates test whether a model can figure out how to operate a mechanism they are not familiar with. A switch is rendered as a colored circle on a cell, while a gate looks like a blocked cell. A model needs to explore and understand how the switches and the gates are associated and what action needs to be executed at what exact cell in the maze in order to be able to resolve the mechanism and make forward progress in the maze. - Dependency chains test whether models are able to reason about the ordering of actions they need to take to resolve mechanisms in a sequence and make forward progress in the maze. For example - a gate that opens only once a switch is toggled can be behind a door that opens only once its key is carried. - Distractors test error recovery - decoy keys, inactive switches, and dead-end branches cost moves to the model without moving them forward in the maze towards the goal. Upon interacting with these distractors, the models need to understand that they play no role in helping them reach their goal, plan the steps to get back on track, and execute the actions in the right order ## Evaluation Setup Designing the complete evaluation required several decisions regarding the configurations of each parameter of the environment as well as the protocol of a model’s interaction with the environment. In order to finalize these decisions, we ran a set of structured ablation experiments across 540 episodes, 12 different settings, on 15 held out mazes. ### Evaluated Models We evaluated 3 frontier vision-language models: Claude Opus 4.8 with adaptive thinking at the xhigh setting, Kimi K2.6 with thinking, and Qwen 3.6 27B with thinking. Claude and Kimi were queried via the native API, and we hosted Qwen 3.6 in parallel on two A100 80GB Virtual Machines to optimize runtime. ### Ablation-based Evaluation Protocol As seen in Table 1 below, the evaluation protocol sweep holds a baseline set of conditions fixed and varies 12 conditions across 7 different axes, each experiment changing exactly one field. We used 15 mazes that are held out from the 50 that were used for the final evaluation in order to ensure that protocol tuning would not leak into the benchmarking scores. The standard baseline evaluation protocol gives the model a prompt stating the task, the action space which includes actions relative to the direction the agent is facing, and the locations of the mechanisms in the maze, but does not give the model information about how to solve the mechanisms. The observation of the current timestep is provided as an image accompanied by the text description of it. To provide historical context on the model’s progress up until the current timestep, the images of the previous 3 states are provided to the model as a part of the observation. Given all the visual and text context, the model is prompted to produce the immediate next action and each query is a single message with no previous turns from its history included. To guide the model on how to complete a maze, one in-context example is given as reference. | Condition | What it changes | Solves / 45 | | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | | cond\_prompt·standardbaseline | The baseline prompt containing the task instruction, action space, and mechanism positions. This prompt does not mention any rules. | 27 | | baseline\_thinking | Not one of the 7 axes we ran ablations for. This setting is the baseline, but with thinking enabled for the models. | 30 | | cond\_prompt·minimal | The prompt is reduced to the task instruction and the action space. | 27 | | cond\_prompt·verbose | Compared to the standard condition, the prompt is extended with explanations about how the mechanisms work. | 25 | | icl\_zero\_shot | The in-context example is removed from the standard setting. | 25 | | qry\_subgoal | The model plans out actions for multiple steps into the future per turn, instead of deciding a per-step action. | 20 | | ctx\_text\_summary | History of the model's progress up to the current timestep is given as a text summary, in place of the three most recent observation frames. | 17 | | hist\_multiturn | Chat history is accumulated across the previous 3 turns, instead of one self-contained message per query. | 15 | | ctx\_current | History of progress is removed entirely, leaving only the current frame as context for the model to take the next action. | 14 | | act\_cardinal | Egocentric directions in the action space are replaced with absolute cardinal directions (North, South, East, West). | 12 | | qry\_full\_trajectory | The model is prompted to provide all the steps it would take to complete the maze in a single response. | 11 | | obs\_image\_only | Only the image of the environment in its current state is provided as context, without any text information. | 5 | Table 1: Ablation experiment results across 12 conditions, 15 mazes, and 3 models to finalize the parameter configurations and evaluation protocol. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/fig01_solve_heatmap-2.png) Figure 3: Heatmap indicating solve rate of each model given each condition in the ablation experiments that were run to finalize evaluation protocol. As seen in Figure 3, not providing text information about the current timestep’s observation significantly reduces performance. On the other hand, varying the amount of detail in the prompt barely seems to affect model performance. Based on this extensive sweep, the evaluation protocol was finalized with the following settings: - Thinking setting on for all the models, with an equal 64k token budget - Minimal prompt setting since models do similarly performance-wise on the different prompt variations and also incentivize models to explore the environment and reason in order to succeed - Image as the only modality of perception of the current state. This was picked intentionally to make the benchmark more challenging and also incentivize the models to learn about the consequences of their actions in the environment by observing and understanding the difference between 2 states - pre and post action, as often needs to be done in real-world workflows ### Task Pool and the Evaluation Set The initial pool of mazes that we created using a combination of hand-design and LLM-based generation included 214 mazes that were all confirmed solvable. The shortest path required to solve a maze ranged from 9 to 106 executable actions across all the mazes. From this initial pool, we picked a set of 50 mazes such that they were balanced and stratified across 5 path-length bands - 2 mazes requiring 20-25 moves, 12 mazes requiring 26-39 moves, 14 mazes requiring 40-59 moves, 8 mazes requiring 60-79 moves, and 14 mazes requiring more than 80 moves. The grid sizes for the selected maze set are 8x8 for 11 mazes, 10x10 for 19 mazes, and 14x14 for 20 mazes. The mazes are grouped into 3 different categories - Scale, Mechanism, and Distractor. The evaluation set included 3 scale mazes - which are mazes that require only navigation, 37 mechanism mazes - which include keys, doors, switches, gates in a dependency order, and 10 distractor mazes which includes decoys on top of mechanisms - such as a wrong key, a switch that does nothing, or a dead-end branch. When generating the initial pool of mazes we varied grid size, topology, dependency chain pattern, mechanism types, mechanism counts, and distractor types and counts. Every maze instance was confirmed reachable by an exhaustive search under a 500,000-state cap, and the optimal action count is computed with a BFS implementation that calculates the actions required to solve the mechanisms as well. We also did 3 further checks to thoroughly validate each maze: - Mechanism necessity - reports any mechanism whose removal does not affect the solvability of the maze - Chain ordering - confirms that each mechanism is unreachable until the prior mechanism in the dependency ordering has been solved - Distractor safety - confirms that no single distractor interaction can make the task unsolvable ## Results and Analysis ### Solve Rates and Progress ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/phase_h_maze_type_matrix-5.png) Figure 4: Maze solves by maze type for each model (Left) and Mean action progress by maze type for each model (Right) As seen in Figure 4, only 6 out of 150 mazes were solved totally by all 3 models. Claude solved 4 of its 50, Kimi 1, and Qwen 1\. 45 out of the 50 mazes were not solved by any of the models. We also observed that all the solves came on mazes with no switches, no distractor objects, and an optimal path length of 45 steps or lesser. By grid size, we see that models solved 3 out of 33 8x8 mazes, 3 out of 57 10x10 mazes, and none of the 60 14x14 mazes. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/progress_grid_actions-3.png) Figure 5: Action progress made by each model on each maze depicted as a heatmap. Each row is a model, and each column is a maze. Star denotes a solve. Action progress is calculated as one minus the remaining oracle actions from the terminal step of a model in a maze over the total oracle actions as calculated by a BFS implementation for the same maze. Figure 5 shows that none of the 3 models came close to the goal tile without actually solving the maze. The median closest approach is 47.5 executable actions from the goal tile, and the only episodes arriving within the 3 actions of the goal tile are the six solves. ### Difficulty Axes Since we chose and designed a maze as our evaluation environment, several components determine how difficult it is for a model to succeed. How long the optimal path is, the grid size, how many mechanisms need to be cleared and in what order, and how many distractors are present, etc. We tested the various components that we thought made a given maze difficult and computed its influence on the action progress score of the models we evaluated. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/path_length_vs_success-2.png) Figure 6: Success rate of models for each bin of optimal steps required to solve the maze (Top) and action progress rate of models for each bin of optimal steps required to solve the maze (Bottom). We found that optimal path length is the largest influence on the difficulty of a maze when compared to the other parameters that make up the environment. When tested one property at a time against the progress, with the model that ran the episode held constant, path length shows an R² value of 0.171 - more than gates that showed 0.122, switches that showed 0.104, and distractor count that showed 0.030. Each property is tested with one regression that predicts an episode's action progress from that property and from which model ran it. R² in simple terms is the share of the variation in progress that this regression accounts for, 1 − SSres / SStot, where SSres is the total squared gap between actual and predicted progress and SStot is the total squared gap between actual and average progress. An R² of 0.171 means the fit removes about 17% of the spread. We see in Figure 6 that performance falls as the optimal path length required to solve the mazes increases, which agrees with the regression calculation above. When grouped by shortest path, mazes of 30 moves or fewer gave 5 solves in 39 episodes at a mean action progress of 0.287 (95% CI 0.163 to 0.423), 31 to 45 moves gave 1 in 30 at 0.240 (0.164 to 0.333), 46 to 60 gave 0 in 15 at 0.141 (0.089 to 0.205), and 61 or more gave 0 in 66 at 0.070 (0.058 to 0.084). When restricted to the 69 episodes on the mazes with optimal solution path lengths of 45 moves or fewer, the influence of gates rise from an R² value of 0.122 to 0.255 and switches from 0.104 to 0.247\. Meanwhile the influence of path length on this restricted subset falls from 0.171 to 0.012, which shows that when the range of optimal path length is restricted, the mechanisms have a relatively significant impact on the difficulty of the maze. ### Failure Modes Our evaluation runs resulted in 0 solves across the 105 episodes containing a switch-maze mechanism. Across all 3 models, despite 132 toggle actions being issued, only 2 of these actions resulted in flipping a switch. 15 of these toggles resulted in successful door opens, but the remaining 115 resulted in nothing. We found it interesting that a model stood on a live switch in 38 out of the 105 mazes containing the switch-gate mechanisms, which means that the models had no issue getting to the tile with the switch, but they were not able to figure out how the mechanism worked in association with a gate. On the contrary to how the models were confounded by the switch-gate mechanism, we observed that the models produced 15 successful door opens. Moreover, 4 of the 6 total solves came on the key-door mazes. One Kimi solve even required a key, a door, a second key, and a second door in sequence. What separates the key-door mechanism from the switch-gate mechanism is the fact that a frontier model is able to associate a key with a door purely based on visual perception due to its prior knowledge, whereas the functioning of a switch which is represented as a circle on a tile needs to be understood by exploring the environment and trying the various actions on a switch cell to observe how the environment changes. Despite relative success with the key-door mechanism, the models interact with both mechanisms inefficiently. We observed that they issued 220 pickup actions which in turn produced only 27 successful key pickups. Similarly, with the toggle action, the models issued 132 of them, of which only 15 actually opened a door, and 2 flipped a switch. Models reach for interaction verbs far more than there is anything for them to interact with in the current state. The switch-gate mechanism failure mode was pretty consistent across all the 3 models. However, we also observed that the models displayed a variety of other failure modes in their evaluations as well. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/failure_mode_frequency.png) Figure 7: Failure mode frequency across episodes for each model ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/step_at_failure_heatmap-2.png) Figure 8: Heatmap displaying the frequency of the step count at failure for each model As seen in Figure 7 and Figure 8, Claude fails early in most of its episodes. Additionally, we saw in the output traces of our results that in 91% of Claude’s last 30 actions before it fails an episode, it pushes forward into a wall. In 41 of its 46 stalled episodes, wall collisions make up for more than half of the final 30 actions. As the steps in a given episode increase, the share of actions leading to wall collision increases, while successful moves fall. In Kimi’s traces we observed that it retreads visited ground quite often. The revisits climb from 2.7% of its actions in the first quarter of steps in an episode to 21.2% in its last. For Qwen, the percentage of moves that were turns climbs from 44.5% to a peak of 59.6% in the third quarter of its steps in an episode, while the share of steps reaching an unseen tile collapses from 24.5% to 2.4% indicating that it produces a lot of unfruitful turn actions as it progresses in the maze. ### Test-time Compute ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/08/tokens_per_new_tile.png) Figure 9: Number of output tokens per unique tile discovered, per episode, shown as a boxplot on a log axis with medians labelled As we can see in Figure 9, Kimi K2.6 and Qwen 3.6 27B reason significantly more and utilize more than 20x the amount of output tokens as Claude to make more progress in the maze. However, as we saw in the Results section, Kimi and Qwen only have 1 solve each compared to Claude’s 4\. This is an indication that more reasoning and tokens spent may increase the proportion of the maze explored by Kimi and Qwen, but does not help with the baseline success rate. Despite all 3 models being given the same output token budget of 64k, we see a clear disparity in the amount of tokens spent across the models. Another observation we made as we looked into the model traces was that Claude’s thinking contracts as it makes progress in an episode. Interestingly, the thinking contracts further when it produces a move that achieves no progress in the maze. Median thinking tokens after a move that achieved nothing is 234 across 1301 turns, whereas it is 579 across 299 turns following a move that worked. ## Discussion **Difficulty of 2D mazes:** The results and analysis from our evaluation in this work showcase that 2D mazes with mechanisms that require causal reasoning to solve, are a very challenging environment for highly capable VLMs. By increasing just one parameter of the environment - optimal path length to solve a maze, the task becomes nearly impossible for frontier models to solve. The difficulty of this underlying substrate will only increase as the mechanisms and distractors are made more challenging - we see this when we restrict the mazes to a certain path length and compute the influence of mechanisms such as switches and gates. Since we believe that this evaluation substrate is a a highly simplified representation of real-world workflows that requires models to take actions over several steps, while encountering novel scenarios, and learning how a new interface must be operated, we think that there is a lot of room for improvement in performance of models and agents engaging in long-horizon workflows. **Inability to learn from exploration:** We also see that frontier VLMs lack the ability to explore, understand, and learn from the environment: As we see in the case of the switch-gate mechanism, models are not able to causally reason in order to associate a toggle action on a cell containing a switch with the gate opening and unlocking more of the maze. This is an important finding as models and agents utilized in real-world workflows constantly come up against interfaces, objects, and scenarios that they have never seen before. Being able to figure out how something needs to be operated without any prior knowledge is an important trait to improve the reliability and capability frontier of a model. **Protection from contamination and saturation:** Due to the nature of the task and environment, we were able to isolate parameters of the environment to understand the effects of each one of them on the difficulty of the maze in the form of structured ablations. The ability to turn any of these knobs to increase difficulty ensures that our benchmark stays free of saturation, and due to the fact that each maze is created as a combination of hand-design and LLM generation - it is entirely new data, thus protecting our benchmark from contamination-related issues. With MultiNet v2.0, we aim to take a step towards solving these 2 important issues that plague most benchmarks that exist today. ## Limitations **Inability of the benchmark to rank the frontier models:** Since we only see 6 out of 150 solves, with the ranking inverting when we choose action progress rate as the main metric, the evaluation results on this version of the benchmark are not sufficient to differentiate between the capabilities of the frontier models **The discovery finding relies on a single mechanism:** Currently, we conclude that the inability of the models to understand and operate switch-gate mechanisms, while they are significantly better at solving key-door mechanisms, indicates their inability to successfully operate a mechanism that does not rely on prior knowledge. However, this claim could be made significantly stronger with more mechanisms that do not rely on prior knowledge. **Visual history limited to 3 previous steps:** In order to implement a minimal harness, understand the true capability of the model in isolation, and maintain a reasonable budget for input token related costs, we restricted the visual history provided to the model in a given timestep to 3 previous frames. In this study, we do not evaluate how the performance of a model changes if the visual history window is increased. **No human baseline:** We did not run an official human baseline to measure the difference between human performance and model performance and truly understand whether the task is easy for humans of various backgrounds and age-groups. **No information provided to the model on steps remaining:** We did not run an ablation to observe whether models tend to perform better when they are told how many steps they have available at a given point, thus providing them with a budget that could enable more cautious movement. ## Conclusion and Next Steps In this release, we build a 2D maze with obstacles as an environment to benchmark the long-horizon action taking and causal reasoning capabilities of frontier vision-language models. We design the environment in a structured manner by running ablation experiments to thoroughly understand the influence of each parameter that makes up the environment. While models performed really poorly on the final evaluation set, we learned a lot about the different failure modes of each of these models. We learned how frontier models find it hard to figure out how to operate a mechanism that they are not familiar with, and also observed how each model had its own way of failing this benchmark. While Claude constantly jammed into walls, Kimi kept visiting previously seen portions of the maze, and Qwen wandered aimlessly through the maze barely getting close to the goal tile. The results also showed how extended reasoning and a large output token budget does not help improve the performance of these models in tasks such as this one. The motivation behind this benchmark was to build an environment representative of real-world workflows where models and agents have to navigate novel interfaces and scenarios to achieve goals that require several steps aided by reasoning and error recovery. While this version of the benchmark already gives us a good idea of the capability frontier, we aim to assess a more realistic version of workflows in the coming versions of the benchmark - those involving multiple domains. Due to the projectable nature of the underlying substrate in this task, we can build an identical environment and task in various domains such as 3D simulation and pure language. Evaluating a frontier model across all projected domains, gives us a quantified understanding of how well models generalize to new interfaces, modes of perception, and action spaces, which is highly essential to solve and automate high-value trajectories in the real world. ## Work with us At Fig, we are building the control layer for AI: systems that perceive an environment, act in it reliably, and improve from the experience. Perception is largely solved. Agency is the open problem: acting dependably in environments nobody explained, recovering from mistakes, and compounding progress over many steps. As this report shows, it will not come from scaling reasoning budget. Measuring progress towards this goal requires evaluation setups and benchmarks representative of complex, real-world workflows. This is what we aim to do with MultiNet, in collaboration with our friends at Manifold Research, MIT, Georgia Tech, and Tufts. If you build models or agents, or work on benchmarking and evaluation, we want to hear from you - whether that means getting your model on our MultiNet v2.0 benchmark, contributing to the environments we are building, or working with us on what comes after. Please fill out [this form](https://app.notion.com/p/3bf4b1d3c487800596bbe4a150962cc0?pvs=21&ref=fig.inc) to get involved, and we will follow up. ## Citation Please cite this work as: ``` Guruprasad, P., Rivera, S., Lu, H., Jain, A., Ren, H. and Sikka, H. (2026) MultiNet 2.0 Preview: Goal Progress Decays with Task Horizon for Frontier VLMs in Interactive 2D Environments. Available at: https://www.fig.inc/multinet-v2-preview ``` Or use the BibTex citation ``` @online{multinet_v2_preview_technical_report_2026, title = {MultiNet 2.0 Preview: Goal Progress Decays with Task Horizon for Frontier VLMs in Interactive 2D Environments}, author = {Pranav Guruprasad and Sean Rivera and Helen Lu and Arushi Jain and Hangliang Ren and Harshvardhan Sikka}, year = {2026}, url = {www.fig.inc/multinet-v2-preview}, note = {MultiNet 2.0 Preview: Interactive 2D Mazes} } ``` ## Acknowledgements We thank Victor Barres for feedback on the design of this environment and benchmark, and for his insights on the evaluation space more broadly. We also thank Greg Kamradt and Yuansheng Ni for early comments on the direction and framing of this work. ## References 1. Y. Zhu et al., ["Establishing Best Practices for Building Rigorous Agentic Benchmarks,"](https://arxiv.org/abs/2507.02825?ref=fig.inc) Jul. 03, 2025, arXiv. doi: 10.48550/arXiv.2507.02825. 2. P. Anderson et al., ["On Evaluation of Embodied Navigation Agents," ](https://arxiv.org/pdf/1807.06757?ref=fig.inc)Jul. 18, 2018, arXiv. doi: 10.48550/arXiv.1807.06757. 3. M. Chevalier-Boisvert et al., ["Minigrid & Miniworld: Modular & Customizable Reinforcement Learning Environments for Goal-Oriented Tasks," ](https://arxiv.org/abs/2306.13831?ref=fig.inc)Jun. 24, 2023, arXiv. doi: 10.48550/arXiv.2306.13831. 4. F. Chollet, ["On the Measure of Intelligence,"](https://arxiv.org/abs/1911.01547?ref=fig.inc) Nov. 05, 2019, arXiv. doi: 10.48550/arXiv.1911.01547. ### Fixing Failures in Browser-Use Models: Why More Data Isn't Enough URL: https://www.fig.inc/blog/fixing-failures-in-browser-use/ Last updated: 2026-06-25T15:07:13.000Z **[Yangyue Wang](https://locke0.github.io/?ref=fig.inc)1, 2**, **[Harshvardhan Sikka](https://www.harshsikka.com/?ref=fig.inc)1, 2**, **[Yash Mathur](https://scholar.google.com/citations?user=TbK0aCoAAAAJ&hl=en&ref=fig.inc)\*2**, **[Tony Zhou](https://www.linkedin.com/in/tony-y-zhou/?utm%5Fsource=share%5Fvia&utm%5Fcontent=profile&utm%5Fmedium=member%5Fios)\*2**, **[Jinu Nyachhyon](https://jinunyachhyon.github.io/?ref=fig.inc)\*2**, **[Pranav Guruprasad](https://www.linkedin.com/in/pranav-guruprasad-82697514a/?ref=fig.inc)1, 2** \* Equal contributions. 1[Fig](https://fig.inc/?ref=fig.inc); 2[Manifold Research Group](https://www.manifoldrg.com/?ref=fig.inc). 0:00 /0:08 1× Relevant links · [Models](https://huggingface.co/figai/UI-TARS-1.5-7B-GUI-Perturbed?ref=fig.inc) · [Dataset](https://huggingface.co/datasets/figai/GUI-Perturbed?ref=fig.inc) · [Paper](https://arxiv.org/abs/2604.14262?ref=fig.inc) · [Demo](https://huggingface.co/spaces/figai/GUI-Perturbed-Finetuned-Result-Viewer?ref=fig.inc) · [Code](https://github.com/ManifoldRG/GUI-DR?ref=fig.inc) · [Cite this](#cite) TL;DR - We ran three LoRA fine-tuning experiments: varying perturbation type, data scale, and real vs. synthetic sources. - Counter-intuitively, augmentation degrades performance rather than improving it. We find this points to issues in model representations and standard fine-tuning methodologies instead of the data itself - We introduce the fine-tuned 7B GUI model trained on GUI-DR generated data to study the effects of synthetic data on the model's GUI grounding capability in supervised post-training. Key Sections · [GUI model skill gaps](#guimodelskillgaps) · [Experimental setup](#experimentalsetup) · [Three experiments, three surprises](#threeexperiments) · [Discussion](#discussion) · [What's next](#whatsnext) GUI Perturbation — Research Series [ Part 1 · Previous report Dataset Release & Data Augmentation Pipeline GUI grounding failures under controlled UI perturbations. Data, tooling, and evaluation protocol. ](https://www.fig.inc/blog/domain-randomization-for-computer-control/) [ Part 2 · Previous report Baseline Evaluations How leading CUA models perform across perturbation types. Structured failure analysis. ](https://www.fig.inc/blog/gui-pertubed-breaking-browser-use-models/) Part 3 · This report Finetuning Experiments Training on perturbation-augmented data. How does finetuning on training data generated via perturbation affect model failure modes. ## Browser-Use & Computer Control Have Cognitive Behavior Gaps The reflex for an unreliable computer-use agent (CUA) is to write a better prompt. Agent Skills, folders of instructions, scripts, and resources that an agent can discover and call, have made that approach both more capable and more popular \[1\]. The premise is reasonable: give the agent better instructions and it should behave better. Prompting cannot supply a behavior the model never learned. Consider booking a flight. Without spatial-relation reasoning, the agent cannot tell whether seat 14A or 14C is the window seat. Without multi-region visual reading, it books May 21 instead of June 21 because it pulled the wrong cell from a dense calendar. Without instruction-ambiguity reasoning, it books the first flight in the list rather than asking which one you meant. Without self-reflection, it follows the wrong checkout flow all the way to the end. Without the ability to refute a premise, it loops forever hunting a menu item that no longer exists, or carries out a dangerous action because it was told to. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/03/image.png) Figure 1: [Sample 119 of 390](https://huggingface.co/spaces/figai/GUI-Perturbed-Data-Viewer?ref=fig.inc), "Click on the button above 'June 19 2023'" The limitations described above are training data problems, not prompting problems. A model picks up the behaviors needed to handle real software only when those behaviors appear in its training data. This post asks one question: can we train these behaviors into a model using GUI-Perturbed data? We find that the obvious approaches fail, and that the way they fail is the useful part. ## Evaluation Gaps to Training Gaps In [Part 2](https://www.fig.inc/blog/gui-pertubed-breaking-browser-use-models/) of this investigation, we found that state-of-the-art GUI models degrade sharply under two conditions: small visual perturbations, and instructions phrased as spatial relations. These models had seen millions of GUI screenshots, yet a change in zoom or a request for "the button above X" was enough to break them. The cause is visible in how CUA training data is usually organized. Standard recipes sort data by surface category: platform, action type, application, UI element type \[3-5, 7\], and try to maximize diversity along those axes. The gaps Part 2 exposed do not lie on those axes. They are gaps in cognitive behavioral coverage: spatial reasoning, instruction disambiguation, invariance to visual appearance. A dataset can be exhaustive across platforms and applications and still contain almost no examples that demand reasoning about where one element sits relative to another. 0:00 /0:05 1× Figure 2: Failure modes identified in [part 2](https://blog.fig.inc/measuring-brittleness-in-gui-grounding-models-using-gui-perturbed?ref=fig.inc) vs. training interventions That points to a direct test: ***If the gaps are behavioral, can we build training data that targets the missing behaviors and fills them?*** As a first step, we study how synthetic grounding data, generated to exercise exactly these behaviors, affects a state-of-the-art model. ## Why GUI Training Data is Hard to Get Right ### Collection is Expensive & Synthesis is Fragile There are two ways to get more grounding data, and each has a characteristic failure mode. - **Real trajectories are expensive.** Collecting real interaction traces at scale is costly. OpenCUA \[6\] and the UI-TARS \[2\] pipeline show what is achievable, but the cost per trajectory stays high and the datasets stay narrow in behavioral diversity. - **Synthetic data is fragile.** Generating data synthetically is the obvious alternative, and it brings its own risk. The Jedi dataset is the cautionary case: synthetic trajectories can look plausible while encoding shortcuts and rendering artifacts that do not transfer to real use, which is why a usable training mix still needs a large fraction of real screenshots \[7\]. × | Synthetic Element | Synthetic Icon | | ------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------- | | ![Original Variant](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/03/jedi-example-1.png) | ![Style Variant](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/03/jedi-example-2.png) | Figure 3: [Jedi](https://osworld-grounding.github.io/?ref=fig.inc) dataset examples. Click on each image to enlarge. The result is that practitioners reach for whatever data is available and hope scale compensates for any distribution mismatch. The experiments below test whether that bet pays off. ### LoRA as the Practical Post-Training Tool Full fine-tuning is impractical at 7B+ parameters, so most teams reach for LoRA (low-rank adaptation): it is fast, memory-efficient, and easy to iterate on \[8\]. LoRA freezes the base weights and learns a small low-rank update on top, and that low-rank constraint is the catch. The rank sets a ceiling on how much representational change the update can express \[9\], and GUI spatial reasoning may demand exactly the deep visual-spatial realignment a low-rank update cannot reach. The DoRA \[10\], GLAD \[11\], and EvoCUA \[12\] results all point the same way: LoRA fine-tuning of vision-language models (VLMs) can degrade capabilities in ways that are hard to predict in advance. ### "Agentic Cognitive Behaviors" By an agentic cognitive behavior we mean something specific: not an action type (click, scroll, type), but a pattern over the full (instruction, observation, thought, action) tuple. It is how a model reasons about and acts on the relational, functional, and visual properties of several on-screen regions at once, whether the screen is a static screenshot or a live application. The behaviors from the opening, spatial reasoning, self-reflection, refutation, and instruction disambiguation, are specific gaps that follow from a distributional limitation. A model that never saw an example requiring spatial-relation reasoning will not acquire it, no matter how many screenshots it has seen. These behaviors live in the interaction between the instruction, the reasoning trace, and the visual input. They cannot be read off any one channel alone, which is why adding more of the same kind of data does not produce them. ## Experimental Setup ### Model: UI-TARS-1.5-7B with LoRA We fine-tune UI-TARS-1.5-7B because it comes from the same model family we evaluated in Part 2 \[13\], allowing us to directly connect findings to training interventions. We use a conservative LoRA configuration: rank 8, representing 0.042% of the model's trainable parameters which lets us test whether lightweight adaptation is sufficient for the representational shifts that GUI grounding requires. ### Training Data We build two training sets so we can hold scale fixed and vary only the kind of data: synthetic and targeted, against real and diverse. 1. **GUI-Perturbed (synthetic, targeted).** We run the Part 1 perturbation pipeline, GUI-DR, over the Mind2Web \[16\] training set, then filter the output with Holo2-30B-A3B \[15\], the current ScreenSpot-Pro \[17\] state of the art at 66.1% accuracy. The split covers four perturbation types (style, text shrink, precision, and an all-combined mix) for 24,935 steps in total, summarized in table 1. | Data Split | Variant Composition | Sample Size | | -------------------------- | ------------------------------- | ----------- | | 6.5k style | style | 6500 | | 6.5k text shrink precision | text shrink + precision | 6500 | | 6.5k all | style + text shrink + precision | 6500 | | 25k all | style + text shrink + precision | 24935 | Table 1: GUI-Perturbed training data splits and their variant compositions. 1. **Salesforce GUI grounding mix (real, diverse).** As a real-data baseline at matched scale, we sample 25k examples uniformly from the Salesforce GUI grounding dataset \[14\], which aggregates several open-source sources (table 2). | Source Dataset | License | | -------------- | -------------------------------- | | Aria-UI | Apache License 2.0 | | OmniAct | MIT License | | Widget Caption | Creative Commons Attribution 4.0 | | UI-Vision | MIT License | | OS-Atlas | Apache License 2.0 | Table 2: Salesforce GUI grounding dataset sources The two data sets allow us to fairly compare synthetic targeted data (GUI-Perturbed) against real diverse data (Salesforce mix) at matched scale. Both experiment 2 & 3 are evaluated on GUI-Perturbed and ScreenSpot-v2 \[21\]. ## Three experiments, three surprises ### Experiment 1: Which Kinds of Perturbations Help? Our first experiment compares augmentation variants to understand which types of perturbation data are most the most impactful on improving grounding. We train separate models on style-only perturbations, on text-shrink-and-precision perturbations, and on the full combined set. The result is counterintuitive. All augmentations lead to slight degradation with text shrink precision only variant resulting in slightly more degradation on average as seen in figure 4\. The most degradation (\~3.3% with direct instruction and no reasoning) is seen on the text shrink variant in GUI-Perturbed eval set. One might expect text shrink perturbations to be the gentlest form of augmentation, changing text size and layout zoom level while preserving everything else. Instead, they produce the largest drop in grounding performance. Figure 4: Baseline vs model variants finetuned on three 6.5k data mixes (mixed style+text shrink + precision / style / text shrink + precision) hit accuracy on GUI-Perturbed ### Experiment 2: Does More Data Help? If targeted data helps even a little, more of it should help more, or at least do no harm. The second experiment scales the training set from 6.5k to 25k samples to test whether more perturbation data improves performance. The standard expectation is that more data improves performance, or at worst plateaus it. Figure 5: Baseline vs 6.5k Mixed vs 25k Mixed hit accuracy on GUI-Perturbed Figure 6: Baseline vs 6.5k Mixed vs 25k Mixed hit accuracy on ScreenSpot v2 We observe amplified degradation due to scaling as seen in figures 5 and 6\. More perturbed data widened the gap from baseline rather than closing it. This contradicts standard scaling intuitions and points to two interacting problems. First, catastrophic forgetting: the distribution shift introduced by perturbed data compounds as the training set grows, pushing the model further from its original capabilities. Second, the LoRA configuration memorizes noise from realistic perturbations instead of learning the invariances the perturbations were designed to teach. The low-rank constraint means the model has limited capacity for new representations, and it spends that capacity fitting artifacts rather than extracting generalizable patterns. ### Experiment 3: Real Data vs Synthetic Data If the problem with synthetic data is a distribution mismatch with real screens, real data should do better. The third experiment tests that directly, comparing the Salesforce mix (real, diverse, drawn from many open-source sets) against GUI-Perturbed (synthetic, targeted at specific perturbations) at the same scale. Figure 7: Baseline vs finetuned variants on 25k Mind2Web perturbed vs 25k Salesforce hit accuracy on GUI-Perturbed Figure 8: Baseline vs 25k Salesforce vs 25k Mind2Web Perturbed hit accuracy on ScreenSpot v2 As seen in figures 7 and 8, neither data set improve performance. Real diverse data degraded the model along different axes than synthetic perturbations, but both degraded it. That points away from the data and toward the recipe or something more fundamental about the model itself. Simple finetuning recipes cannot make the representational change GUI spatial reasoning needs: the model has to alter how it maps visual patches to spatial meaning, and that is a deeper change than a small percentage of its parameters can carry. ## Discussion ### GUI Models are More Sensitive to Data Distribution than Data Scale The standard intuition in machine learning is that more diverse data leads to better generalization. What we observe with LoRA SFT on GUI grounding tasks is different: data scale and diversity matter less than distribution alignment with the target capability. Small amounts of misaligned data cause disproportionate degradation because the low-rank update has limited capacity and allocates it to fitting whatever signal is strongest in the training distribution, even if that signal is noise. This has practical implications. Practitioners who collect or generate more data without carefully controlling its distributional properties may find that their models get worse, not better. Scale is not a substitute for alignment. ### LoRA SFT is Insufficient for Visual-Spatial Alignment GUI grounding requires shifting how the model relates visual patches to spatial semantics, a representational change at the model's feature level, not a behavioral adjusTRent that can be addressed with a LoRA. The findings are consistent with results from the DoRA paper on LoRA sensitivity, the analysis of fine-tuning representation shift for multimodal LLMs, and work on conditional mixture of LoRA approaches. Cross-entropy loss alone may also be insufficient for grounding alignment. The loss optimizes next-token prediction over the action output, but it does not directly supervise the spatial reasoning that produces the correct action. A model can learn to produce plausible-looking coordinate outputs without improving its internal spatial representations. Our [baseline evaluation](https://blog.fig.inc/measuring-brittleness-in-gui-grounding-models-using-gui-perturbed?ref=fig.inc) provides additional evidence. UI-TARS-1.5, trained on Qwen2.5VL-7B likely through further SFT and/or RL on CUA trajectory data (training details not public), achieves worse relational accuracy (35.0%) than the base Qwen2.5-VL (45.0%), despite improving on direct grounding. GTA1, which adds GRPO with step-level click reward on top of UI-TARS-1.5, recovers to 65.8%. The progression suggests that trajectory level supervised fine-tuning on GUI trajectories can improve direct element matching while degrading spatial reasoning, and that reinforcement learning with step-level grounding-specific reward is more effective at teaching geometric understanding. ### Current Benchmarks Mask these Dynamics Perhaps the most concerning finding is that without perturbation-based evaluation, we would not have detected these degradation patterns. Models that score well on fixed-scene benchmarks can degrade under training interventions that are designed to help them. If we had evaluated performance using only on standard benchmarks, we potentially would have arrived at a different conclusion.might have concluded that the training worked, or at least that it was harmless. GUI-Perturbed as an evaluation tool is essential for honest measurement of training interventions. This reinforces our belief that perturbation-based data is not just useful for stress-testing models, it is necessary for understanding whether training is making progress on the capabilities that matter. ## Scope and Limitations **Training method coverage.** We evaluate LoRA at a single rank configuration. Full fine-tuning, higher-rank LoRA, QLoRA \[18\], and RL-based post-training (such as GRPO \[14\]) are all plausible alternatives that may yield different results. Our findings apply to the conservative LoRA regime that most practitioners use, but they should not be read as a general claim about all post-training methods. **Data coverage.** We compare two data sources at matched scale. Broader augmentation strategies, curriculum-based approaches, and combinations of real and synthetic data remain unexplored. ## What's Next ### Behavior-Driven Data Curation Today's CUA training data is organized by surface features: platform, application, element type. Our results argue for organizing it by the behaviors it teaches instead, visual reasoning, error correction, refutation, clarification, and spatial-relation reasoning. Many of these behaviors are moving targets. They change with software updates, vary across tasks, and differ between users, and they show up through interaction rather than static annotation. Scaling behavioral coverage will likely take new curation methods: paraphrasing instructions for variety, auto-annotating interaction traces from user feedback, and building pipelines that ground training data in a desired distribution over instruction, observation, and behavior. ### Better Post-Training Recipes LoRA SFT with cross-entropy loss is not enough on its own. The directions we find most promising combine stages and signals: multi-stage training that pairs SFT with RL (as in SpatialLadder \[19\] and GuirlVG \[20\]), higher-rank adaptation that gives the model more room for representational change, and process reward models that supervise grounding decisions step by step rather than scoring a whole sequence at once. ### Richer Learning Signals from Environment State Current GUI training operates on a simple mapping: (screenshot, instruction) produces an action. What is missing is a representation of the next state, the result of taking that action. Without next-state information, the model has no way to learn from the consequences of its actions during training. Better computer state representations could unlock richer credit assignment and more efficient learning signals. This connects back to the domain randomization thesis from [Part 1](https://www.fig.inc/blog/domain-randomization-for-computer-control/): just as robotic policies benefit from simulators that provide full state feedback, GUI agents could benefit from environment representations that go beyond static screenshots. Building those representations is a direction we are actively exploring at Fig. ## Conclusion Across this series, [developed a new domain randomization approach for GUI data](https://www.fig.inc/blog/domain-randomization-for-computer-control/), [used it to expose systematic weaknesses in state-of-the-art models](https://www.fig.inc/blog/gui-pertubed-breaking-browser-use-models/), and, in this work, attempted to fix those weaknesses through training. While the training results negatively impacted model performance, they were very informative. We learned that naive data augmentation with conservative fine-tuning does not close the behavioral gaps we identified. Style perturbations degrade rather than improve, more data amplifies the effect, and real data vs. synthetic data both fail when the training recipe cannot support the representational changes the task requires. The path forward requires rethinking both training data coverage and how models learn from it. On the data side, we need to evolve from surface-level diversity (more platforms, more applications) to behavioral diversity (more reasoning patterns, more failure recovery, more spatial understanding). On the training side, we need recipes that go beyond LoRA SFT: higher-capacity adaptation, reinforcement learning from grounding feedback, and learning signals that capture the consequences of actions rather than just the actions themselves. ## Work With Us At Fig, we are building the control layer for AI: systems that perceive an environment, act in it reliably, and improve from the experience. Perception is largely solved. Agency is the open problem: acting dependably in real environments, recovering from mistakes, and compounding over time. It will not come from scaling language models, and as this post shows, it will not come from bolting a prompt or a light fine-tune onto a model that never learned the behavior in the first place. [Subscribe](https://www.fig.inc/blog/gui-pertubed-breaking-browser-use-models/#/portal/signup) to stay updated as we make progress! If you'd like to work on the next frontier of intelligent systems, [reach out](mailto:contact@fig.inc)! ## Citation Please cite this work as follows: ```latex @online{training_on_gui_perturbed_technical_report_2026, title = {Fixing Failures in Browser-Use Models: Why More Data Isn't Enough}, author = {Yangyue Wang and Harshvardhan Sikka and Yash Mathur and Tony Zhou and Jinu Nyachhyon and Pranav Guruprasad}, year = {2026}, url = {www.fig.inc/blog/fixing-failures-in-browser-use/}, note = {Part 3: Finetuning Experiments} } ``` ## References \[1\] "[Overview](https://agentskills.io/home?ref=fig.inc)," Agent Skills. Accessed: Mar. 10, 2026. \[2\] Y. Qin et al., "[UI-TARS: Pioneering Automated GUI Interaction with Native Agents](https://arxiv.org/abs/2501.12326v1?ref=fig.inc)," arXiv.org. Accessed: Mar. 10, 2026. \[3\] J. Mu et al., "[GUI-360°: A Comprehensive Dataset and Benchmark for Computer-Using Agents](https://arxiv.org/abs/2511.04307?ref=fig.inc)," Nov. 10, 2025, arXiv. doi: 10.48550/arXiv.2511.04307. \[4\] H. Li, J. Chen, J. Su, Y. Chen, Q. Li, and Z. Zhang, "[AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs](https://arxiv.org/abs/2502.01977?ref=fig.inc)," Jun. 07, 2025, arXiv. doi: 10.48550/arXiv.2502.01977. \[5\] S. Nayak et al., "[UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction](https://arxiv.org/abs/2503.15661?ref=fig.inc)," May 06, 2025, arXiv. doi: 10.48550/arXiv.2503.15661. \[6\] X. Wang et al., "[OpenCUA: Open Foundations for Computer-Use Agents](https://arxiv.org/abs/2508.09123?ref=fig.inc)," Oct. 04, 2025, arXiv. doi: 10.48550/arXiv.2508.09123. \[7\] T. Xie et al., "[Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis](https://arxiv.org/abs/2505.13227?ref=fig.inc)," Oct. 24, 2025, arXiv. doi: 10.48550/arXiv.2505.13227. \[8\] E. J. Hu et al., "[LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685?ref=fig.inc)," Oct. 16, 2021, arXiv. doi: 10.48550/arXiv.2106.09685. \[9\] G. Pantazopoulos and E. B. Özyiğit, "[An Efficient Training Pipeline for Reasoning Graphical User Interface Agents](https://arxiv.org/abs/2511.08172?ref=fig.inc)," Nov. 14, 2025, arXiv. doi: 10.48550/arXiv.2511.08172. \[10\] S.-Y. Liu et al., "[DoRA: Weight-Decomposed Low-Rank Adaptation](https://arxiv.org/abs/2402.09353?ref=fig.inc)," Jul. 09, 2024, arXiv. doi: 10.48550/arXiv.2402.09353. \[11\] Y. Peng, P. Wang, J. Liu, and S. Chen, "[GLAD: Generalizable Tuning for Vision-Language Models](https://arxiv.org/abs/2507.13089?ref=fig.inc)," Jul. 17, 2025, arXiv. doi: 10.48550/arXiv.2507.13089. \[12\] T. Xue et al., "[EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience](https://arxiv.org/abs/2601.15876?ref=fig.inc)," Jan. 23, 2026, arXiv. doi: 10.48550/arXiv.2601.15876. \[13\] "[ByteDance-Seed/UI-TARS-1.5-7B](https://huggingface.co/ByteDance-Seed/UI-TARS-1.5-7B?ref=fig.inc)," Hugging Face. Accessed: Mar. 10, 2026. \[14\] Y. Yang et al., "[GTA1: GUI Test-time Scaling Agent](https://arxiv.org/abs/2507.05791?ref=fig.inc)," Oct. 03, 2025, arXiv. doi: 10.48550/arXiv.2507.05791. \[15\] "[Hcompany/Holo2-30B-A3B](https://huggingface.co/Hcompany/Holo2-30B-A3B?ref=fig.inc)," Hugging Face. Accessed: Mar. 10, 2026. \[16\] X. Deng et al., "[Mind2Web: Towards a Generalist Agent for the Web](https://arxiv.org/abs/2306.06070?ref=fig.inc)," Dec. 09, 2023, arXiv. doi: 10.48550/arXiv.2306.06070. \[17\] K. Li et al., "[ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use](https://arxiv.org/abs/2504.07981?ref=fig.inc)," Apr. 04, 2025, arXiv. doi: 10.48550/arXiv.2504.07981. \[18\] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, "[QLoRA: Efficient Finetuning of Quantized LLMs](https://arxiv.org/abs/2305.14314?ref=fig.inc)," May 23, 2023, arXiv. doi: 10.48550/arXiv.2305.14314. \[19\] H. Li et al., "[SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models](https://arxiv.org/abs/2510.08531?ref=fig.inc)," Oct. 09, 2025, arXiv. doi: 10.48550/arXiv.2510.08531. \[20\] W. Kang, B. Lei, G. Liu, C. Ding, and Y. Yan, "[GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning](https://arxiv.org/abs/2508.04389?ref=fig.inc)," Aug. 06, 2025, arXiv. doi: 10.48550/arXiv.2508.04389. \[21\] K. Cheng *et al.*, “[SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents](https://arxiv.org/abs/2401.10935?ref=fig.inc),” Feb. 23, 2024, *arXiv*: arXiv:2401.10935\. doi: 10.48550/arXiv.2401.10935. ### GUI-Perturbed: Breaking Browser-use Models using Domain Randomization URL: https://www.fig.inc/blog/gui-pertubed-breaking-browser-use-models/ Last updated: 2026-06-25T15:13:06.000Z **[Yangyue Wang](https://locke0.github.io/?ref=fig.inc)1, 2**, **[Harshvardhan Sikka](https://www.harshsikka.com/?ref=fig.inc)1, 2**, **[Yash Mathur](https://scholar.google.com/citations?user=TbK0aCoAAAAJ&hl=en&ref=fig.inc)\*2**, **[Tony Zhou](https://www.linkedin.com/in/tony-y-zhou/?utm%5Fsource=share%5Fvia&utm%5Fcontent=profile&utm%5Fmedium=member%5Fios)\*2**, **[Jinu Nyachhyon](https://jinunyachhyon.github.io/?ref=fig.inc)\*2**, **[Pranav Guruprasad](https://www.linkedin.com/in/pranav-guruprasad-82697514a/?ref=fig.inc)1, 2** \* Equal contributions. 1[Fig](https://fig.inc/?ref=fig.inc); 2[Manifold Research Group](https://www.manifoldrg.com/?ref=fig.inc). 0:00 /0:05 1× --- TL;DR - We introduce a baseline study of 7B GUI models using **GUI-Perturbed** to stress test CUA models and understand *what* agents fail at and *why*. - A detailed failure mode analysis showcases common GUI failure modes from spatial reasoning, false visual heuristics to CoT reasoning's effect. Key Sections · [How do GUI models fail](#howmodelsfail) · [The triple alignment problem](#thetriplealignment) · [Experimental setup](#experimentalsetup) · [Results](#results) · [Discussion](#discussion) · [Model failure modes](#failuremodeexamples) Relevant links · [Code](https://github.com/ManifoldRG/GUI-DR?ref=fig.inc) · [Dataset](https://huggingface.co/datasets/figai/GUI-Perturbed?ref=fig.inc) · [Results Viewer](https://huggingface.co/spaces/figai/GUI-Perturbed-Baseline-Result-Viewer?logs=build&ref=fig.inc) · [Cite this](#cite) GUI Perturbation — Research Series [ Part 1 · Previous report Data Augmentation Pipeline GUI grounding failures under controlled UI perturbations. Data, tooling, and evaluation protocol. ](https://www.fig.inc/blog/domain-randomization-for-computer-control/) [ Part 2 · This report Dataset Release & Baseline Evaluations How leading CUA models perform across perturbation types. Structured failure analysis. ](https://www.fig.inc/blog/gui-pertubed-breaking-browser-use-models/) [ Part 3 · Last Report Fine-tuning Experiments & Model Checkpoint Training on perturbation-augmented data. How does fine-tuning on training data generated via perturbation affect model failure modes. ](https://www.fig.inc/blog/fixing-failures-in-browser-use/) GUI models scoring above 90% on ScreenSpot-v2 fail to find the target element [when you set web page zoom to 70%](#failureat0%5F7zoom) \[1\]. Same website, same layout, same UI elements. Just smaller. These models were trained on hundreds of thousands to millions of GUI screenshots through supervised fine-tuning and reinforcement learning stages yet they still cannot adapt to a change in zoom. [ ![Original variant](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/part-2-flight-example-orig.png) ](#v1) Original Model Result ✓ Correct [](#fmc-close)[✕](#fmc-close)![Original variant](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/part-2-flight-example-orig.png) [ ![Precision variant at 70% zoom](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/part-2-flight-example-precision.png) ](#v2) Precision Variant (70% zoom) Model Result ✗ Clicked on fake 'View Deal' ad [](#fmc-close)[✕](#fmc-close)![Precision variant at 70% zoom](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/part-2-flight-example-precision.png) Figure 1: [Sample: 21 of 390](https://huggingface.co/spaces/figai/GUI-Perturbed-Data-Viewer?ref=fig.inc): "Click on 'View Deal' button for flight '#2125, #2126'", Direct Instruction, No Reasoning. Result: UI-TARS1.5-7B clicked on the fake 'view deal' button in the ads after the 70% zoom. GUI grounding looks solved. On fixed-scene benchmarks like ScreenSpot-v2, 7B models now score above 90%, and it is tempting to read those numbers as evidence that perception is no longer the bottleneck for computer-use agents (CUAs). The numbers are real, but they are measured on screens that never move. Real users zoom, restyle, and resize, and production websites are redesigned constantly. The question the industry should care about is not how well a model does on a frozen screenshot, but how much of that performance survives contact with ordinary variation. So we ask one question: **how much of a GUI grounding model's benchmark accuracy is stable under perturbation, and how much of it is memorized?** Because the three models we study share a base checkpoint but differ in post-training, we can ask a sharper version too: ***does each additional stage of GUI-specialized post-training buy real robustness, or does it only raise the fixed-scene score?*** Our goal here is to separate those two things. In [Part 1](https://www.fig.inc/blog/domain-randomization-for-computer-control/), we introduced [GUI-DR](https://github.com/ManifoldRG/GUI-DR?ref=fig.inc), a data augmentation pipeline that varies visual scenes and instructions along controlled axes to stress test CUA model's GUI grounding capability. In this post, we use it to create a [dataset composed of visual scene and the instruction variations](https://huggingface.co/datasets/figai/GUI-Perturbed?ref=fig.inc) created along controlled axes. This dataset serves as a benchmark to evaluate three state-of-the-art models that share the same base checkpoint but differ in their post-training recipes, and we report where they break. We find that: - **Visual perturbations degrade models that benchmarks call production-ready.** A change as small as setting browser zoom to 70% drops accuracy by 2 to 6 points across all three models. - **Spatial relational instructions are the weakest point.** Asking for "the button above X" instead of "the submit button" costs 27 to 56 points, the largest single effect we see. - **Reasoning is not uniformly good.** A chain of thought helps on hard relational tasks and hurts on easy direct ones, and a model post-trained for direct coordinate prediction is harmed by it everywhere. - **More GUI-specialized post-training does not fix any of this.** The same weaknesses persist from the base model through two further stages of GUI training. ## The Triple Alignment Problem Grounding a GUI instruction is harder than it looks, because the model has to align three different things at once and a benchmark score collapses all three into one number. - **Visual alignment*: identifying an element's appearance in pixel space, its shape, color, size, and boundaries.* - **Functional alignment*: knowing what the element does, telling an input field from a display label or a clickable button from a static icon.* - **Geometric alignment*: resolving spatial relationships between elements, "above," "next to," "the one between X and Y."* ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/Screenshot-2026-03-05-at-4.29.25---PM-1.png) Figure 2: Triple alignment in GUI agent perception Most benchmarks test the three entangled together \[2, 3\], so when a model fails the score cannot say which alignment broke. GUI-Perturbed is built to stress visual and geometric alignment independently, which lets us attribute a failure to one axis rather than guess. (Our current perturbations do not isolate functional alignment the way AutoGUI's instructions do \[4\]; we leave that for future work.) This framing carries through the rest of the post. For every result, we name the alignment axis it implicates. ### Problem Formulation A computer-use agent acts over many steps: it sees a screen, picks an action, the screen changes, and it sees the next one. That full setting is a partially observable Markov decision process (POMDP) \[9-11\], and recent agentic models are trained against it \[12, 13\]. We do not need the full machinery here, because our evaluation isolates a single step. **The POMDP tuple** defines the CUA problem structure: hidden app states, observations (screenshots + goal), actions, transition dynamics, and reward \[12\]. \\\[ \\mathcal{M} \\;=\\; \\langle\\, \\mathcal{S},\\; \\mathcal{A},\\; \\mathcal{O},\\; \\underbrace{\\mathcal{T}(s\_{t+1}\\mid s\_t, a\_t)}\_{\\text{transition}},\\; \\mathcal{R} \\,\\rangle \\\] At each step, a CUA model receives an instruction I and an observation OO (the screenshot), optionally produces a chain of thought T, and outputs an action A. The full step can be written as: \\\[ O\_t \\;\\sim\\; \\mathcal{Z}(O\_t \\mid s\_t), \\qquad (t\_t,\\, a\_t) \\;=\\; \\mathrm{VLM}\_\\theta\\!\\left( I,\\; O\_{1:t},\\; t\_{1:t-1},\\; a\_{1:t-1} \\right), \\quad I \\in \\mathcal{I} \\\] **Triple alignment** is the key representational challenge: to select a correct action at at , the model must be able to interpret Ot along three axes visual appearance, geometric relation, functional affordance of the GUI from screenshots simultaneously. \\\[ a\_t \\;=\\; \\pi\_\\theta\\!\\left(I,\\; O\_{1:t},\\; t\_{1:t-1},\\; a\_{1:t-1} \\right), \\qquad O\_t \\;\\supseteq\\; \\left(\\, \\underbrace{O\_t^{\\mathrm{vis}}}\_{\\substack{\\text{visual}\\\\\\text{appearance}}},\\; \\underbrace{O\_t^{\\mathrm{geo}}}\_{\\substack{\\text{geometric}\\\\\\text{layout}}},\\; \\underbrace{O\_t^{\\mathrm{func}}}\_{\\substack{\\text{functional}\\\\\\text{affordance}}} \\,\\right) \\\] Our evaluation isolates the grounding step: given (I,O), predict the correct element. This removes multi-step dependencies and focuses the evaluation on the single-step alignment problem described above. We do evaluate models in both reasoning and no-reasoning configurations, which introduces a planning-like component through the thought trace T, and we report on its effects below. ## Experimental Setup ### Three Models, One Base Checkpoint ① Pretrain Qwen2.5VL-7B Vision-Language Model 4.1T token pretraining Builds on ② Specialized Fine-Tune UI-TARS1.5-7B GUI-Specialized Trajectories \~50B GUI-focused tokens Builds on ③ GUI GRPO Training GTA1-7B Grounding Agent + o3 Planner Test-time scaling Figure 3: Model lineage diagram We study three 7B models built on the same base weights, so that any difference in robustness is attributable to post-training rather than to scale or architecture: - **Qwen2.5-VL-7B*, the base vision-language model (VLM). It saw some GUI trajectories during long-context pre-training but had no dedicated GUI fine-tuning \[5\].* - **UI-TARS-1.5-7B*, initialized from Qwen2.5-VL and trained further on end-to-end GUI trajectories \[8\]. The exact recipe is not public; the UI-TARS papers \[6\] point to supervised fine-tuning (SFT) and reinforcement learning.* - **GTA1-7B*, initialized from UI-TARS-1.5 and trained further on GUI data with step-level Group Relative Policy Optimization (GRPO) \[7\].* | | Qwen2.5VL-7B | UI-TARS1.5-7B | GTA1-7B | | ------------- | -------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | Architecture | | | | | Base model | Qwen2.5VL-7B | Qwen2.5VL-7B | UI-TARS1.5-7B (grounding) \+ o3 (planner) | | Training | | | | | Stages | Visual Pre-Training (1.5T tokens) Multimodal Pre-Training (2T tokens) Long-Context Pre-Training (0.6T tokens) \+ SFT + DPO | Continual Pre-training (GUI knowledge) Annealing Phase (SFT) DPO Phase Based on UI-TARS recipe, UI-TARS1.5 training recipe not public \[7\] | RL Optimization: GRPO (Group Relative Policy Optimization) Click reward mechanism | | Training Data | | | | | Volume | 4.1T tokens | $\\sim$50B tokens | — | | Sources | Interleaved image-text VQA Image captions & OCR Visual knowledge Video grounding Document parsing Agent interaction data | 18.4M grounding elements (web, mobile, desktop) 6M GUI tutorials 151.4k action traces Reflective online traces | Aria-UI \[15\] OmniAct \[16\] Widget Caption \[17\] UI-Vision \[18\] OS-Atlas \[19\] (lightly cleaned) | Table 1: Model architecture & training comparison We chose this lineage deliberately, the models share the same architecture and same base weights with different training recipes. Any performance variation between the three is directly attributable to post-training, not architecture, which allows us to precisely ask: *how does each additional stage of GUI-specialized training affect robustness under perturbation?* ### Stress Test Matrix We evaluate three 7-B models from the same lineage on eight task variants from GUI-Perturbed, enabling comparisons along three axes visual variants, instruction types, and reasoning mode: | Axis | Levels | Cardinality | | ---------------------------------- | ----------------------------------------------------------------------------------------- | ----------- | | Visual Perturbation | \\(\\{\\text{Original},\\ \\text{Style},\\ \\text{Precision},\\ \\text{Text Shrink}\\}\\) | 4 | | Instruction Type | \\(\\{\\text{Directional},\\ \\text{Relational}\\}\\) | 2 | | Reasoning Mode | \\(\\{\\text{With CoT},\\ \\text{Without CoT}\\}\\) | 2 | | **Total Configurations per Model** | **$4 \\times 2 \\times 2 = 16$** | | Table 2: Stress-testing configuration space using [GUI-Perturbed](https://huggingface.co/datasets/figai/GUI-Perturbed?ref=fig.inc) ### Evaluation Metrics We focus our analysis primarily on hit rate, flip rate, and net delta, which measure the models' overall performance and sensitivity under perturbation: - **Hit rate** measures whether the predicted coordinate falls within the bounding box of the target element. We report 95% bootstrap confidence intervals (10,000 resamples) for all hit rates; exact binomial (Clopper--Pearson) intervals agreed within 0.2 pp throughout and are omitted for brevity. **This is our primary metric: a grounding prediction either lands on the right element or it does not.** - **Flip rate** measures the fraction of matched pairs whose binary outcome (hit/miss) changed between the original and the perturbed condition. Each perturbation test uses *n* \= 390 matched sample pairs (same task and step evaluated on both the original and perturbed screenshots). **A high flip rate indicates the model's output is sensitive to the perturbation, regardless of whether accuracy improves or degrades on average.** - **Net Delta** measures the difference in hit rate between the original and perturbed conditions (original minus perturbed), with 95% bootstrap CI. **A positive Delta indicates overall degradation.** We also test significance with McNemar's test, which compares the number of samples that degraded (*b*: correct to incorrect) against those that improved (*c*: incorrect to correct) under perturbation, ignoring samples whose outcome did not change. We report *p*\-values with continuity correction; when the number of discordant pairs (*b* \+ *c*) is below 25, we use the exact binomial test instead. Additionally, MSE, normalized MSE, and normalized distance provide additional signal on error magnitude but did not reveal trends beyond what the primary metrics capture. ## Results ### Visual Perturbations Break Models that Benchmarks Call Robust All three models degrade under visual perturbations, including a change as small as browser zoom. Models scoring above 90% on fixed-scene benchmarks drop on perturbed versions of the same pages. | | | | Flip Rate | Net $\\Delta$ (%) | | | | | | ----------- | ----------- | --------- | --------- | ----------------- | ---------- | ---------- | ------------------------------------- | ---- | | Model | Pert. | Base Acc. | Dir. | Rel. | Dir. | Rel. | $\\boldsymbol{b}$ / $\\boldsymbol{c}$ | Sig. | | GTA-1 | Precision | 79.3 | 10.3% | 21.5% | +3.6\*\* | +7.9\*\*\* | 169/79 | 3/4 | | | Style | | 9.7% | 21.5% | +1.3 | −0.3 | 126/118 | 0/4 | | | Text Shrink | | 4.2% | 16.7% | +0.4 | +1.8 | 90/73 | 0/4 | | Qwen2.5-VL | Precision | 66.0 | 13.1% | 16.4% | +3.3 | +4.9\* | 147/83 | 2/4 | | | Style | | 8.7% | 19.1% | +2.8\*\* | +0.1 | 120/97 | 1/4 | | | Text Shrink | | 7.8% | 14.4% | +0.1 | +2.8 | 98/75 | 0/4 | | UI-TARS-1.5 | Precision | 63.0 | 13.1% | 18.7% | +6.2\*\*\* | +5.4\*\* | 169/79 | 4/4 | | | Style | | 11.2% | 19.2% | +2.4 | +1.0 | 132/105 | 0/4 | | | Text Shrink | | 6.9% | 14.1% | −0.8 | +0.0 | 79/85 | 0/4 | Table 3: Perturbation robustness of baseline models (n = 390 matched sample pairs per test). The zoom and text-size perturbations are worth emphasizing as they are not exotic transformations. Any user adjusting their browser zoom or system font size produces exactly this kind of variation. The fact that models trained on millions of GUI screenshots cannot handle a zoom change suggests they are memorizing absolute spatial positions rather than understanding relational structure. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/Figure-4-3.png) Figure 4: A cross model comparison against hit accuracy. All three models degrade 2-6% on precision variant (70% zoom). UI-TARS1.5 shows the largest drop. In our experiments, all 3 models experience 2-6% degradation on precision variant (70% zoom) compared to the original variant with direct instructions as seen in [figure 4](#fig4). UI-TARS1.5-7B shows the most degradation with precision variant. Specifically, precision perturbation (70% zoom) produced statistically significant accuracy drops in 9 of 12 paired comparisons (McNemar's p < 0.05) as seen in [table 3](#table3). Although, style perturbation reached significance in only 1 of 12, and text-shrink in 0 of 12, the lack of significant net degradation does not mean those perturbations left predictions unchanged. Style perturbation flipped 14.9% of all predictions (698 of 4,680), nearly matching precision's 15.5% flip rate (726 of 4,680). Style flips were roughly symmetric between degraded and improved, so the net effect washed out. This points to a useful distinction between robustness (whether net accuracy holds) and consistency (whether individual predictions stay stable). All three visual perturbation types destabilize predictions substantially; only precision does so in a systematically harmful direction. Some combinations of perturbations unexpectedly improved accuracy on individual configurations (e.g., style on relational+CoT for GTA-1: 63.1% to 65.1%, +2.1 pp), though none reached significance (all p > 0.4). Similar non-significant improvements appeared for text-shrink on UI-TARS-1.5 direct queries (92.8% to 93.8%, p = 0.45). Our hypothesis is that some perturbations inadvertently increase whitespace between elements or enlarge text in ways that make grounding easier. This is a useful diagnostic signal in itself: it tells us which visual properties these models are most sensitive to. ### Spatial Reasoning is the Weakest Link The sharpest performance drops come from relational instruction variants. When the instruction asks the model to identify an element by its spatial relationship to a neighbor (“click the button above X”), performance falls well below what the same models achieve on direct instructions (“click the submit button”). | Benchmark | Qwen2.5VL-7B | UI-TARS1.5-7B | GTA1-7B | | --------------------------- | ------------------- | ------------------- | ------------------- | | ScreenSpot-v2 | 88.8 | 89.7 | 92.4 | | ScreenSpot-Pro | 27.6 | 42.0 | 50.1 | | OSWorld | — | $27.4 \\pm 2.2\\%$ | 45.2 (with o3) | | OSWorld-G | 27.7 | 64.2 | 67.7 | | GP-Unperturbed (Direct) | 86.9 | 91.0 | 92.8 | | GP-Unperturbed (Relational) | 45.0 (↓\\(-41.9\\)) | 35.0 (↓\\(-56.0\\)) | 65.8 (↓\\(-27.1\\)) | Table 4: Comparing our direct and relational query approaches on unperturbed data (GP-Unperturbed, Direct) with the queries of other datasets; **Bold** \= best score per benchmark. This gap between direct and relational performance is consistent across all three models and both reasoning modes ranging from 27.1% to 56.0% as seen in [table 4](#table4). It is the single largest effect we observe in the evaluation, larger than any visual perturbation effect. We also examined whether models exhibit systematic directional biases in their spatial errors. Our directional hit rate analysis shows that models perform unevenly across spatial directions, with consistently higher accuracy on instructions involving the direction “right” compared to other directions. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/Figure-5.png) Figure 5: Hit accuracy by model for relational query directions. Models show consistently higher accuracy on 'right' instructions. Directional bias may reflect training distribution or patchification artifacts. We note that our directional analysis relies partly on manual labeling to correct reference landmark coordinates, which limits the scale at which we can draw conclusions. The directional asymmetry we observe is suggestive of biases introduced during training or by the visual patchification process, but confirming this would require a larger controlled study. ### Reasoning Helps on Hard Tasks, Hurts on Easy Ones The effect of reasoning mode is not uniformly positive. Enabling reasoning (the thought trace TT in our formulation) produces different outcomes depending on task difficulty. On simple direct grounding tasks, reasoning introduces unnecessary deliberation that can actively mislead the final prediction. The model “overthinks” a task that the base visual grounding would handle correctly without intermediate reasoning. On more complex relational tasks, reasoning recovers some performance by providing useful intermediate structure: the model can reason about spatial relationships step by step rather than attempting to resolve them in a single forward pass. ![image.png](https://codahosted.io/docs/zujAXxemDw/blobs/bl-VMYV2juCGl/d6afec19e0e4349b213211cc6ab6d8b3dca7ef20159c3943a50e561dca6741390abda2a29b4edc3c337e57e6c25313a4f90be9044592cc175b67612982c6f2d138f3ab0da16e8b49b663c4339f7da1b8468525da721bc4fe329a77da1f34c96625917589) Figure 6: Instruction: "**Click on 'Notifications' div*". Model output: "**Thought: I noticed that there is a "Notifications" option in the left sidebar, which is exactly what I need to click on. This option is located just below "Privacy and data" and above "Security and logins." By clicking on it, I can access the notification management page. Action: (None, 'Notifications')*" GTA1 provides a particularly instructive case. It was further trained to predict coordinates directly, and it suffers degradation from reasoning on both simple and harder relational tasks. Its post-training has optimized it for direct coordinate prediction, and the reasoning trace interferes with that pipeline regardless of task complexity. The implication is that blanket “enable reasoning” is not the right strategy. Models need exposure to diverse reasoning styles during post-training, and they need to calibrate when to reason and how much. The optimal reasoning style and length likely varies by task. ### Failure Mode Taxonomy We conducted a qualitative analysis of representative failures across all models and configurations. Several recurring patterns emerge. | Failure Mode | Definition | | ------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | | spatialClick Region Error | The model selects the correct UI element conceptually but clicks the wrong physical area of it. | | spatialLocation Hallucination | The model correctly identifies what to click but fabricates or misplaces its on-screen coordinates. | | reasoningReasoning Drift | Adding explicit reasoning causes the model's grounded action prediction to deteriorate. | | semanticGoal Hallucination | The model invents user intentions or interface functionality not specified in the instruction. | | semanticInstruction Misinterpretation | The model misunderstands what the instruction refers to and selects a related but incorrect element. | | visualVisual Confusion | The model relies on superficial visual cues (shape, color, position) and misidentifies the functional element. | | spatialSpatial Reasoning Error | The model incorrectly interprets relative spatial relationships such as left, right, above, or below. | | semanticText Matching Bias | The model interacts with visible text matching the instruction without properly grounding it to the correct UI element. | Table 5: Failure mode definitions The full set of annotated failure examples is available in the [appendix section 2](#failuremodeexamples). ## Discussion ### Models Lack Spatial Relational Understanding ![image.png](https://codahosted.io/docs/zujAXxemDw/blobs/bl-4t1blNR3oz/6bd4b34b564ed8a922d58ebc2982e5eef824f3cd6d500243f10fd5d9d5b04783491e9789677f649bb06e1ee59f54c99584c0a8bb7e7392ba9492a2635d6794ef478fdeed02002d520de3bf8289764cd86949f85499e1f55c850bd068f37009463349af8f) Figure 7: [****Sample 216 of 390**](https://huggingface.co/spaces/figai/GUI-Perturbed-Data-Viewer?ref=fig.inc)****:** “Click on the button above ‘FOURMIDABLE’ img”; UI-TARS1.5-7B prediction: "Thought: I noticed that there is a green heart icon located above the "FOURMIDABLE" logo. According to the task requirements, I need to click on this heart icon. It's positioned right next to the logo, and clicking on it should allow me to complete the task. Action: click(start\_box='(1899,138)')"; UI-TARS1.5 confuses what is on the right with what is above. The sharpest performance drops in our evaluation come from relational instructions. Models can find an element by name but cannot resolve “the field above X.” This is not a data quantity problem. These models have been trained on millions of screenshots. The issue is representational: current architectures do not build structured spatial models of GUI layouts. They encode visual features at the patch level without maintaining an explicit spatial graph of element relationships. This maps directly to the geometric alignment axis from our framing. Visual and functional alignment may be adequate for direct grounding tasks, but geometric alignment, the ability to reason about spatial relationships between elements, is where current models fall short. ### Visual Heuristics are Static and Fragile Models learn fixed visual associations (white rectangle at the top of the screen equals search bar) that break on any layout or style change. In the zoom perturbation results, we see models clicking on advertisement elements that happen to occupy the spatial position where the target element used to be at the original zoom level. The model has memorized a position, not learned a function. In production, websites update their designs regularly. A model relying on static visual heuristics is one deployment away from failure. This is a visual alignment problem: the model’s visual representations are too tightly coupled to the specific pixel-level appearances in the training distribution. ![image.png](https://codahosted.io/docs/zujAXxemDw/blobs/bl-YHqC77JaKm/1a92834a294214dd105d79b2560f193461e23cdedf84f6b1688e9f27f3ed5707a4176692e4c7ab43f5ba533e93d958105c19f3617c9cda89df04af59942585118f880e4c9f00d8e3b431c1e9fadbbbfe28d6aef7f50ee891ae00a5fe3d532a43cef6e071) Figure 8: A new UI version change could render many GUI agents useless on the same website \[14\] ### Reasoning is a Double-Edged Sword for Grounding Pulling the reasoning results together: deliberation helps on hard relational tasks and hurts on easy direct ones, and GTA-1 sharpens the point. A model post-trained for direct coordinate prediction is harmed by reasoning in every condition, because the trace disrupts the pipeline its training optimized for. The implication for post-training is that reasoning should be treated as a learnable, task-conditioned skill rather than a switch to flip on or off. Models need exposure to varied reasoning styles and lengths during training so they can learn when deliberation helps and when it does not. ## Scope and Limitations **Model coverage.** We evaluate three models from one base checkpoint lineage and while this approach isolates the effect of post-training recipes, it may not generalize to models with different base architectures or scales. Broader coverage is future work. **Perturbation coverage.** We evaluate on eight variants from GUI-Perturbed. More perturbation types and combinations are possible, and interactions between perturbation types (ex: zoom combined with relational instructions) remain unexplored. ## What’s Next This report used GUI-Perturbed as a benchmark to uncover systematic weaknesses in state-of-the-art GUI grounding models that persist across all three models despite increasingly specialized post-training. Part 3 will explore the natural next question: whether we can fix these weaknesses with better training data. We use GUI-Perturbed for data augmentation and measure whether targeted training closes the gaps we identified here. At Fig, we're building the control layer for AI: systems that perceive an environment, act in it reliably, and improve from the experience. Frontier models have largely solved perception. Agency is the open problem — acting dependably in real environments, recovering from mistakes, and compounding over time — and it will not come from scaling language models alone. Computer use is where we begin, and grounding is where the gap first surfaces: a model that can read a screen but cannot reliably act on it is not yet in control. GUI-Perturbed measures that gap precisely. Closing it, across software today and every environment over time, is the work ahead. [Subscribe](https://www.fig.inc/blog/domain-randomization-for-computer-control/#/portal/signup) to stay updated as we make progress! If you'd like to get involved, reach out to contact@fig.inc ## Citation Please cite this work as: ```latex @online{measuring_gui_models_robustness_technical_report_2026, title = {GUI-Perturbed: Breaking Browser-use Models using Domain Randomization}, author = {Yangyue Wang and Harshvardhan Sikka and Yash Mathur and Tony Zhou and Jinu Nyachhyon and Pranav Guruprasad}, year = {2026}, url = {www.fig.inc/blog/gui-pertubed-breaking-browser-use-models/}, note = {Part 2: Baseline evaluation} } ``` # References \[1\] K. Cheng et al., ["SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents,"](https://arxiv.org/abs/2401.10935?ref=fig.inc) Feb. 23, 2024, arXiv: arXiv:2401.10935\. doi: 10.48550/arXiv.2401.10935. \[2\] T. Xie et al., ["OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments,"](https://arxiv.org/abs/2404.07972?ref=fig.inc) May 30, 2024, arXiv: arXiv:2404.07972\. doi: 10.48550/arXiv.2404.07972. \[3\] K. Li et al., ["ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use,"](https://arxiv.org/abs/2504.07981?ref=fig.inc) Apr. 04, 2025, arXiv: arXiv:2504.07981\. doi: 10.48550/arXiv.2504.07981. \[4\] H. Li, J. Chen, J. Su, Y. Chen, Q. Li, and Z. Zhang, ["AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs,"](https://arxiv.org/abs/2502.01977?ref=fig.inc) Jun. 07, 2025, arXiv: arXiv:2502.01977\. doi: 10.48550/arXiv.2502.01977. \[5\] S. Bai et al., ["Qwen2.5-VL Technical Report,"](https://arxiv.org/abs/2502.13923?ref=fig.inc) Feb. 19, 2025, arXiv: arXiv:2502.13923\. doi: 10.48550/arXiv.2502.13923. \[6\] Y. Qin et al., ["UI-TARS: Pioneering Automated GUI Interaction with Native Agents,"](https://arxiv.org/abs/2501.12326?ref=fig.inc) Jan. 21, 2025, arXiv: arXiv:2501.12326\. doi: 10.48550/arXiv.2501.12326. \[7\] Y. Yang et al., ["GTA1: GUI Test-time Scaling Agent,"](https://arxiv.org/abs/2507.05791?ref=fig.inc) Oct. 03, 2025, arXiv: arXiv:2507.05791\. doi: 10.48550/arXiv.2507.05791. \[8\] U.-T. Team, ["UI-TARS - Next-generation native GUI agent model,"](https://seed-tars.com/?ref=fig.inc) UI-TARS. Accessed: Mar. 04, 2026. \[9\] M. Lu et al., ["Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs,"](https://arxiv.org/abs/2511.19773?ref=fig.inc) Nov. 24, 2025, arXiv: arXiv:2511.19773\. doi: 10.48550/arXiv.2511.19773. \[10\] L. Zhao et al., ["Seeing is Believing: Belief-Space Planning with Foundation Models as Uncertainty Estimators,"](https://arxiv.org/abs/2504.03245?ref=fig.inc) Apr. 04, 2025, arXiv: arXiv:2504.03245\. doi: 10.48550/arXiv.2504.03245. \[11\] H. Li et al., ["VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators,"](https://arxiv.org/abs/2510.00406?ref=fig.inc) Oct. 01, 2025, arXiv: arXiv:2510.00406\. doi: 10.48550/arXiv.2510.00406. \[12\] J. Wu et al., ["OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks,"](https://arxiv.org/abs/2601.20650?ref=fig.inc) Feb. 02, 2026, arXiv: arXiv:2601.20650\. doi: 10.48550/arXiv.2601.20650. \[13\] K. Team et al., ["Kimi K2.5: Visual Agentic Intelligence,"](https://arxiv.org/abs/2602.02276?ref=fig.inc) Feb. 02, 2026, arXiv: arXiv:2602.02276\. doi: 10.48550/arXiv.2602.02276. \[14\] B. Oliveira and C. Teixeira Lopes, ["The Evolution of Web Search User Interfaces - An Archaeological Analysis of Google Search Engine Result Pages,"](https://dl.acm.org/doi/10.1145/3576840.3578320?ref=fig.inc) in Proceedings of the 2023 Conference on Human Information Interaction and Retrieval, Austin TX USA: ACM, Mar. 2023, pp. 55–68\. doi: 10.1145/3576840.3578320. \[15\] Y. Yang et al., "[Aria-UI: Visual Grounding for GUI Instructions](https://arxiv.org/abs/2412.16256?ref=fig.inc)," 2024, arXiv: arXiv:2412.16256. \[16\] R. Kapoor et al., "[OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web](https://arxiv.org/abs/2402.17553?ref=fig.inc)," 2024, arXiv: arXiv:2402.17553. \[17\] Y. Li et al., "[Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements](https://aclanthology.org/2020.emnlp-main.443/?ref=fig.inc)," in *Proc. EMNLP*, 2020, pp. 5495–5510. \[18\] S. Nayak et al., "[UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction](https://arxiv.org/abs/2503.15661?ref=fig.inc)," 2025, arXiv: arXiv:2503.15661. \[19\] Z. Wu et al., "[OS-ATLAS: A Foundation Action Model for Generalist GUI Agents](https://arxiv.org/abs/2410.23218?ref=fig.inc)," 2024, arXiv: arXiv:2410.23218\. doi: 10.48550/arXiv.2410.23218. ## Appendix ### 1\. Model Performance with MSE, NMSE, and D\_norm (Normalized Distance) ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/Figure-9.png) Figure 9: Cross-model bounding box center MSE across perturbation conditions. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/Figure-10.png) Figure 10: Cross-model normalized MSE (NMSE) across perturbation conditions. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/Figure-11.png) Figure 11: Cross-model Euclidean distance from predicted point to bounding box center, normalised by box diagonal. Values above 1.0 indicate predictions outside the target box. ### 2\. Failure Mode Qualitative Examples [![Click Region Error](https://codahosted.io/docs/zujAXxemDw/blobs/bl-5TGnHcA2mX/36b447bc25deaf00e620368eb7747285a2f3b82d893571d6fa5fc80d08f2e4856bcfa8e999d71ed6c37b910ca6a8e9c1e85b1b9f5646473ee47819faf692b999da609dad329a14561a816bdc525313f48c0043b045ef5d9d25526a1d0c92a9db7d27267b)](#c1) Click Region Error Instruction Click on 'Done' buttonAction: click(start\_box='(639,438)') Model misidentifies clicking the area next to 'Done' as equivalent to clicking the 'Done' text itself. [](#fmc-close)[✕](#fmc-close)![Click Region Error](https://codahosted.io/docs/zujAXxemDw/blobs/bl-5TGnHcA2mX/36b447bc25deaf00e620368eb7747285a2f3b82d893571d6fa5fc80d08f2e4856bcfa8e999d71ed6c37b910ca6a8e9c1e85b1b9f5646473ee47819faf692b999da609dad329a14561a816bdc525313f48c0043b045ef5d9d25526a1d0c92a9db7d27267b) [![Location Hallucination](https://codahosted.io/docs/zujAXxemDw/blobs/bl-VMYV2juCGl/d6afec19e0e4349b213211cc6ab6d8b3dca7ef20159c3943a50e561dca6741390abda2a29b4edc3c337e57e6c25313a4f90be9044592cc175b67612982c6f2d138f3ab0da16e8b49b663c4339f7da1b8468525da721bc4fe329a77da1f34c96625917589)](#c2) Location Hallucination Instruction & Model Prediction Click on 'Notifications' divThought: I noticed that there is a "Notifications" option in the left sidebar… This option is located just below "Privacy and data" and above "Security and logins." Action: (None, 'Notifications') Model correctly locates 'Notifications' in its reasoning but clicks a completely wrong coordinate. [](#fmc-close)[✕](#fmc-close)![Location Hallucination](https://codahosted.io/docs/zujAXxemDw/blobs/bl-VMYV2juCGl/d6afec19e0e4349b213211cc6ab6d8b3dca7ef20159c3943a50e561dca6741390abda2a29b4edc3c337e57e6c25313a4f90be9044592cc175b67612982c6f2d138f3ab0da16e8b49b663c4339f7da1b8468525da721bc4fe329a77da1f34c96625917589) [![Reasoning Drift](https://codahosted.io/docs/zujAXxemDw/blobs/bl-PYhfTvXkmY/3fa1f1e04eee71e83bbe290f4552957257a67f5841a13b5313ae1cd24a534ecac3865e0fb74453d46bf8693175b3b64fc5b599211deae2412598ca507b730fd654941a91d87130b19a0b2785a98bafc1198c682542a4d18d55b442fe6390a2dea5090883)](#c3) Reasoning Drift Instruction & Model Prediction Click on 'Verizon' imgThought: The task is to click on the 'Verizon' image… The image is located at the bottom of the product details section. Action: {"name": "computer\_use", "arguments": {"action": "left\_click", "coordinate": \[1245, 768\]}} "Located at the bottom" in the model's own reasoning misleads it to click an unrelated image at the bottom of the screen. [](#fmc-close)[✕](#fmc-close)![Reasoning Drift](https://codahosted.io/docs/zujAXxemDw/blobs/bl-PYhfTvXkmY/3fa1f1e04eee71e83bbe290f4552957257a67f5841a13b5313ae1cd24a534ecac3865e0fb74453d46bf8693175b3b64fc5b599211deae2412598ca507b730fd654941a91d87130b19a0b2785a98bafc1198c682542a4d18d55b442fe6390a2dea5090883) [![Goal Hallucination](https://codahosted.io/docs/zujAXxemDw/blobs/bl-oOJ2Ls2_E4/55dd526f8b45bfe4753b0b963d58aefe36132f566b04761bb7ab775f877705439a8cf32c4250e6dc0029f259b828a6baea92cf49fd76b343a1159a334607360befff1ddaf7cdcf020de76cc9f65b54f293e1612c28fee7267ae459cc0523e22b0cae2aee)](#c4) Goal Hallucination Instruction & Model Prediction Click on the button above 'May 17th 2023'Thought: I noticed a gray heart-shaped button in the upper right corner… clicking on it should allow me to save this cooking activity to my favorites. Action: (1416,317) Model invents a "save to favorites" goal absent from the instruction, then clicks accordingly. [](#fmc-close)[✕](#fmc-close)![Goal Hallucination](https://codahosted.io/docs/zujAXxemDw/blobs/bl-oOJ2Ls2_E4/55dd526f8b45bfe4753b0b963d58aefe36132f566b04761bb7ab775f877705439a8cf32c4250e6dc0029f259b828a6baea92cf49fd76b343a1159a334607360befff1ddaf7cdcf020de76cc9f65b54f293e1612c28fee7267ae459cc0523e22b0cae2aee) [![Instruction Misinterpretation](https://codahosted.io/docs/zujAXxemDw/blobs/bl-vDFl32r5XP/e1af3363804f859003f9d5db10596f6cfa33f5237c55f1f8babc8313dba0e7acb995286cc3d1c738e2c3504f804ab463f8979bd159becfaa0367064778b3f28b6074946fe45f0595635d8fd68741dbe6b2fc090523b88b905b909cfaf564382abbc391dd)](#c5) Instruction Misinterpretation Instruction & Model Prediction Click on 'Egypt'Thought: There is a dropdown menu… To view trending content in Egypt, click on the dropdown button to the right of "Egypt." Action: click(start\_box='(620,488)') Model reinterprets "click on 'Egypt'" as "click the dropdown arrow beside Egypt." [](#fmc-close)[✕](#fmc-close)![Instruction Misinterpretation](https://codahosted.io/docs/zujAXxemDw/blobs/bl-vDFl32r5XP/e1af3363804f859003f9d5db10596f6cfa33f5237c55f1f8babc8313dba0e7acb995286cc3d1c738e2c3504f804ab463f8979bd159becfaa0367064778b3f28b6074946fe45f0595635d8fd68741dbe6b2fc090523b88b905b909cfaf564382abbc391dd) [![Visual Confusion](https://codahosted.io/docs/zujAXxemDw/blobs/bl-yro6LSbqd_/0a8927541f2a7fcefd2cf6ecc7c8f7218aab4e926450eabefc5dc8b198026365b799d9420402c84259043ed0e691c43ff682f218906b70b0d558d0941e60a725bb9ce2545f229e99e27df5d275ea0f32755236c29a32a14f281aff298c17e85344bf7bc3)](#c6) Visual Confusion Instruction & Model Prediction Type 'bed sheets queen' in 'Search: suggestions appear below' searchboxThought: There are two search boxes at the top… The one on the left is the one I need. Action: click(start\_box='(314,28)') Model mistakes a light-colored button with faint text for the target search box. [](#fmc-close)[✕](#fmc-close)![Visual Confusion](https://codahosted.io/docs/zujAXxemDw/blobs/bl-yro6LSbqd_/0a8927541f2a7fcefd2cf6ecc7c8f7218aab4e926450eabefc5dc8b198026365b799d9420402c84259043ed0e691c43ff682f218906b70b0d558d0941e60a725bb9ce2545f229e99e27df5d275ea0f32755236c29a32a14f281aff298c17e85344bf7bc3) [![Spatial Reasoning Error](https://codahosted.io/docs/zujAXxemDw/blobs/bl-Oi1LEmPf7l/162d01a641ad9f597c9cbb95f607187e9c226720e7e587627294019af36df4f7b5a3c17cedca8a5eab52004f63b5b09801e54f7b2c24b9b1aea437c585d1e73560d6380746bac3985fbb2afa2450494873d3f5a73c855afb236dc3e0f6d572b589d5a2e5)](#c7) Spatial Reasoning Error Instruction & Model Prediction Click on the link to the left of 'Side effects'Thought: There is a link "Interactions" located to the left of "Side effects"… positioned in the middle of the navigation bar. Action: (998,517) Model correctly names the target link but clicks a button on the right instead of the left. [](#fmc-close)[✕](#fmc-close)![Spatial Reasoning Error](https://codahosted.io/docs/zujAXxemDw/blobs/bl-Oi1LEmPf7l/162d01a641ad9f597c9cbb95f607187e9c226720e7e587627294019af36df4f7b5a3c17cedca8a5eab52004f63b5b09801e54f7b2c24b9b1aea437c585d1e73560d6380746bac3985fbb2afa2450494873d3f5a73c855afb236dc3e0f6d572b589d5a2e5) [![Text Matching Bias](https://codahosted.io/docs/zujAXxemDw/blobs/bl-1g3ulUFgIn/5c8c82953fe059bd3adb85b19f9545c9e5221c02689a00f333a47840d775792cc92519281443c2746f4e26945558dc99dd736d49001d91f577c642aff2abdc15a58b470dc3dcacf4714967b5cbfbbee4b90f6c57be7a496df2a6815344a6eeb46e2c2ab4)](#c8) Text Matching Bias Instruction Click on 'First Name' textboxAction: click(start\_box='(1242,509)') Model clicks the "First Name" label text rather than the input field beneath it. [](#fmc-close)[✕](#fmc-close)![Text Matching Bias](https://codahosted.io/docs/zujAXxemDw/blobs/bl-1g3ulUFgIn/5c8c82953fe059bd3adb85b19f9545c9e5221c02689a00f333a47840d775792cc92519281443c2746f4e26945558dc99dd736d49001d91f577c642aff2abdc15a58b470dc3dcacf4714967b5cbfbbee4b90f6c57be7a496df2a6815344a6eeb46e2c2ab4) ### Domain Randomization for Computer Control URL: https://www.fig.inc/blog/domain-randomization-for-computer-control/ Last updated: 2026-07-04T23:35:53.000Z **[Yangyue Wang](https://locke0.github.io/?ref=fig.inc)1, 2**, **[Harshvardhan Sikka](https://www.harshsikka.com/?ref=fig.inc)1, 2**, **[Yash Mathur](https://scholar.google.com/citations?user=TbK0aCoAAAAJ&hl=en&ref=fig.inc)\*2**, **[Tony Zhou](https://www.linkedin.com/in/tony-y-zhou/?utm%5Fsource=share%5Fvia&utm%5Fcontent=profile&utm%5Fmedium=member%5Fios)\*2**, **[Jinu Nyachhyon](https://jinunyachhyon.github.io/?ref=fig.inc)\*2**, **[Pranav Guruprasad](https://www.linkedin.com/in/pranav-guruprasad-82697514a/?ref=fig.inc)1, 2** \* Equal contributions. 1[Fig](https://fig.inc/?ref=fig.inc); 2[Manifold Research Group](https://www.manifoldrg.com/?ref=fig.inc). 0:00 /0:09 1× GUI-DR restyles, repositions, and removes DOM elements on real webpages --- TL;DR - GUI models scoring 90%+ on standard benchmarks fail under basic visual variations like a 70% browser zoom. Current benchmarks can't detect this because they evaluate on fixed scenes with fixed instructions. We quantify how far performance drops in Part 2\. [Subscribe](https://www.fig.inc/#/portal/signup) to get it in your inbox. - Our stress-testing framework for GUI grounding applies domain randomization from robotics, varying visual scenes and instructions along controlled axes to expose fragile model behaviors. - We introduce GUI-DR, an open-source data augmentation pipeline for generating perturbation variants from real web pages. Key Sections · [The White Rectangle Problem](#whiterectangle) · [Building GUI-Perturbed](#buildingguiperturbed) · [Dataset at a Glance](#datasetataglance) · [Get involved](#getinvolved) Relevant links · [Code](https://github.com/ManifoldRG/GUI-DR?ref=fig.inc) · [Cite this](#cite) ## The White Rectangle Problem Modern GUI grounding models can locate a “Submit” button with high precision, identify form fields from natural-language instructions, and navigate complex web interfaces. Yet they confuse a browser's search bar with the formula bar in Google Sheets. Both are white rectangles near the top of the screen. Mistakes like these are the demo-to-production gap that keeps GUI models stuck in the lab. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/image-1.png) Figure 1: OpenAI's Operator confuses the browser search bar with the Google Sheets formula bar mid-task. Both are white rectangles near the top of the screen. This is a systematic failure: models ground to visual primitives like shape, position, and color rather than functional semantics \[17\]. A white rectangle at the top of the screen represents “text input,” regardless of whether it is a search bar, a formula bar, or a URL field. The model has skewed representation of what the element might do. Current evaluation datasets can't tell us how widespread the white rectangle problem is \[1,3-12\]. They evaluate on fixed scenes with fixed instructions: a specific screenshot, a referring expression, a single correct answer. That measures peak performance under curated conditions, not how models degrade when layout, zoom, or wording shift, which is much closer to production. The question is whether we can measure grounding robustness systematically: *Instead of only measuring peak accuracy on a fixed scene, can we measure how models hold up as scenes and instructions vary?* In this technical report, we introduce GUI-Perturbed, a dataset built on domain randomization principles that varies visual scenes and instructions along controlled axes to expose fragile grounding. We describe the dataset, the perturbation methodology, and the design decisions behind it. ## Fixed Scenes Hide Fragile Models Existing computer-using agent (CUA) evaluation datasets share a common structure: a fixed screenshot, a fixed instruction, and a fixed ground-truth target. Benchmarks like OSWorld \[3\], ScreenSpot-v2 \[5\], ScreenSpot-Pro \[6\], and OSWorld-G \[4\] each contribute valuable coverage of specific scenarios and applications. But they all evaluate under the same assumption: that the test set’s visual scene and instruction distribution is representative of real world scenarios. In production, this assumption breaks constantly. Websites ship new themes. Browser zoom levels vary across users. Dark mode inverts color relationships. Users describe the same element in different ways depending on context. A model that scores 90% on a fixed test set may score far lower once any of these variables shift. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/03/Screenshot-2026-03-01-at-3.08.30---PM.png) Figure 2: GUI agent dataset comparison \[1,3-12\]. Scene variability: Fixed = no variation; Live = uncontrolled real-world changes; Perturbed = controlled variation. GUI-Perturbed† is web-only; cross-platform is left for future work. What we need is evaluation data that varies these conditions systematically, so we can measure robustness, not only peak performance. For this we borrow a technique from robotics: domain randomization. GUI Perturbation — Research Series [ Part 1 · This report Data Augmentation Pipeline GUI grounding failures under controlled UI perturbations. Data, tooling, and evaluation protocol. ](https://www.fig.inc/blog/domain-randomization-for-computer-control/) [ Part 2 · Next Report Dataset Release & Baseline Evaluations How leading CUA models perform across perturbation types. Structured failure analysis. ](https://www.fig.inc/blog/gui-pertubed-breaking-browser-use-models/) [ Part 3 · Last Report Fine-tuning Experiments & Model Checkpoint Training on perturbation-augmented data. How does fine-tuning on training data generated via perturbation affect model failure modes. ](https://www.fig.inc/blog/fixing-failures-in-browser-use/) ## Sim-to-Real to Demo-to-Production Domain randomization is a standard technique for bridging the gap between simulation and the real world \[13\]. During training, we randomize visual properties of the simulator (textures, lighting, object colors, camera angles) so the policy is forced to learn features that are invariant to surface-level variation. A robot that has seen a red cup, a blue cup, and a transparent cup in training is more likely to generalize to a cup it has never seen than one trained on a single appearance. The benefits are well-established. Domain randomization forces invariance to irrelevant visual features \[16\]. It exposes failure modes that fixed test sets miss. And it scales to large numbers of scenarios without manual curation: you generate new training or evaluation data by sampling new random variations. The parallel to GUI agents is direct. Models trained and evaluated on fixed screenshots are analogous to policies trained in a single simulator skin. They memorize the visual shortcuts of their training distribution (where elements tend to appear, what colors they tend to be, which shapes correlate with which functions) instead of learning the structural relationships between elements. When the skin changes, the policy breaks. Robotics ![Domain randomization in robotics](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/02/DR-robotics.webp) GUI ![Domain randomization in GUI environments](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/06/GUI-Domain-Randomization.png) Figure 3: Domain randomization in robotics vs. GUI environments \[14\] Applying domain randomization to GUIs, however, is a different engineering problem. Robotic simulators provide programmatic control over every visual parameter: change a texture map, adjust a light source, swap out an object mesh \[15\]. GUI environments do not offer similar interface parameters. Changing the appearance of a desktop application typically requires application-specific integration, and most production software exposes limited visual controllability. Our workaround is to operate on MHTML archives of real web pages. MHTML files capture a complete snapshot of a rendered web page, including HTML, CSS, images, and layout, in a single archive. They also preserve the DOM (Document Object Model) structure, which gives us programmatic access to the same elements a browser renders visually. We can add, remove, restyle, and reposition actual DOM elements rather than being limited to pixel-level image transforms. Think of the MHTML file as our simulator. With it, we can randomize the visual environment while keeping the underlying page structure intact. ## Building GUI-Perturbed Building GUI-Perturbed comes down to three choices: which sub-problem to evaluate, what to use as a controllable simulator, and how to perturb it. We isolate step-level grounding, treat Mind2Web's MHTML archives as our simulator, and perturb along two axes: the visual scene and the instruction. ### Isolating Step-Level Grounding We focus on a single, well-defined sub-problem: given a screenshot and a natural language instruction referring to a specific GUI element, can the model correctly identify that element? We deliberately exclude planning, navigation, and multi-step execution so we can attribute failures to grounding rather than upstream errors. If a model fails a multi-step task, it is hard to tell whether the failure came from misreading the instruction, locating the element, or choosing the action. By isolating grounding, we get clean signal. In the language of our domain randomization analogy, this is single-move evaluation: grading each grounding decision independently, the way an analysis engine grades individual chess moves rather than judging by the outcome of the whole game. ### Mind2Web as Our Simulation Engine We build GUI-Perturbed on top of the Mind2Web dataset, which provides MHTML archives of real websites alongside annotated interaction traces \[10\]. Each MHTML file captures a complete web page that we can load, manipulate, and re-render. This gives us a key advantage over screenshot-only approaches. With raw screenshots, perturbation options are limited to pixel-level operations: color shifts, crops, rotations, noise injection. With DOM access, we can make semantically meaningful changes: restyle a button, reposition a form field, swap the order of navigation items, change the theme of the entire page. These are the kinds of variations that occur naturally in production and that fixed-scene benchmarks miss. ### Two Axes of Perturbation A grounding model takes two inputs: a visual scene (the screenshot) and an instruction (the natural language description of the target element). We perturb both. 0:00 /0:10 1× Figure 4: Two-axis perturbations: visual scene axis × instruction axis **Visual scene perturbations** change the rendered page while preserving the target element. The goal is to alter the visual context (neighboring elements, page style, layout properties) so that a model relying on visual shortcuts will fail while a model with structural understanding will succeed. **Instruction perturbations** change how the target element is described. The same button can be referred to as “the submit button,” “the green button at the bottom of the form,” or “the button below the email field.” Each phrasing requires different capabilities: keyword matching, visual attribute recognition, or spatial reasoning. ![Screenshot 2026-02-23 at 1.47.48 PM.png](https://codahosted.io/docs/zujAXxemDw/blobs/bl-WDSEllQ7hG/dc2ec5569f1ae7789fb5383c956154547c8f1f43960344784e4ecfa8e96a9fd055b47744f1b0d7cce1b26d48eff49c8090cff4e9ddda7c79367050aa11d0aee35d18fefedc79155bc2c55c36c961fc54ea696b4c13981b4ef064561ac1a213646bc8f3b6) Figure 5: GUI-Perturbed data generation algorithm Returning to our domain randomization analogy: visual perturbations are like changing the simulator’s textures and lighting conditions. Instruction perturbations are like giving the robot a different way of specifying the goal. Grounding has to survive both. ### Relational Instructions Half of our instruction perturbations use relational instructions: referring expressions that identify the target element by its spatial or functional relationship to other elements on the page, rather than by the target’s own properties. For example: - “Click on ‘unread message’ above the ‘reservation email’” - “Click on the arrow icon under the second image to expand the comments” Compare these to direct instructions like “click the blue submit button” or “click the search icon.” Direct instructions require the model to match a description to a single element. Relational instructions require the model to identify a reference landmark, reason about a spatial relationship (above, below, next to, between), and then locate the target relative to that landmark. We define relational instructions precisely: a relational instruction is one that identifies the target element for an action based on a given reference landmark and direction descriptions. This distinction matters for two reasons. First, relational instructions reflect how humans actually refer to GUI elements in practice. When guiding someone through a UI over the phone, we say “click the button next to the search bar,” not “click the element at coordinates (450, 230).” Second, relational instructions interact with visual perturbations in diagnostic ways. If we move a neighboring element, does the model still resolve “next to the search bar” correctly? This creates a natural test of whether the model maintains a structured spatial representation of the page or relies on memorized co-occurrence patterns. The term relational instruction carries across all three parts of this series: it is central to the evaluation results in Part 2 and the training experiments in Part 3. ## Anatomy of a Perturbation × | Original Variant | Style Variant | | ------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------- | | ![Original Variant](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/02/anatomy-of-perturbation-1.webp) | ![Style Variant](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/02/anatomy-of-perturbation-2.webp) | Figure 6: Original variant vs. style variant examples from GUI-Perturbed. Click on each image to enlarge. The figure above shows an example directly from GUI-Perturbed. On the left is the original Mind2Web screenshot with its associated instruction. On the right is a perturbed version of the same page. A robust grounding model should recognize that despite the visual changes, the target element still serves the same function and still satisfies the instruction. Each perturbation is designed so that a model relying on surface-level visual associations (element position, surrounding colors, layout proximity) should fail it, while a model that understands the element's functional role should succeed. ## Dataset At A Glance × | Variant | N | Description | | ----------- | --- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | Original | 390 | Obtained the screenshots of the pages directly rendered from Mind2web mhtml files (the original pages) | | Style | 390 | Obtained the screenshots after injecting the original pages with templated CSS and JS code to randomize their button orders and element styles | | Precision | 390 | Obtained the screenshots after scaling the pages to 0.7 | | Text Shrink | 390 | Obtained the screenshots after scaling down the text font size | | Original | Style | Precision | Text Shrink | | ----------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | | ![Original](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/02/original.webp) | ![Style](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/02/style.webp) | ![Precision](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/02/precision.webp) | ![Text Shrink](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2026/03/part2textshrinkdatasetataglance.jpg) | Figure 7: Perturbation Examples. Click on each image to enlarge. ## Scope and Limitations **Perturbation realism.** Not all perturbations produce pages that look like production websites. We prioritize diagnostic coverage over photo-realism. A perturbation that no real website would produce can still reveal a meaningful model weakness: if a model fails when we change the background color of a page, that failure tells us something about the model’s reliance on color as a grounding cue, regardless of whether the specific color is realistic. **Instruction diversity.** People refer to GUI elements in many ways. Our instruction perturbations cover a useful subset of referring expressions but not the full distribution of natural language. Expanding this coverage, particularly for colloquial and ambiguous references, is a direction for future work. **Web domain only.** This release covers web-based GUIs. Desktop applications, mobile interfaces, and cross-application workflows present different challenges and are out of scope for this release. ## What’s Next At Fig, we're building the control layer for AI: systems that perceive an environment, act in it reliably, and improve from the experience. Frontier models have largely solved perception. Agency is the open problem — acting dependably in real environments, recovering from mistakes, and compounding over time — and it will not come from scaling language models alone. Computer use is where we begin, and grounding is where the gap first surfaces: a model that can read a screen but cannot reliably act on it is not yet in control. GUI-Perturbed measures that gap precisely. Closing it, across software today and every environment over time, is the work ahead. [Subscribe](https://www.fig.inc/p/b22a18b0-1433-424a-8656-dc098ff8cbb4/?member%5Fstatus=free#/portal/signup) to our newsletter to follow our progress. Aside from its practical use as a benchmark, building GUI-Perturbed pushed us toward deeper questions about grounding itself: how models represent interface elements, when they fall back on visual shortcuts instead of function, and how spatial and relational reasoning hold up as a scene changes. Domain randomization gives us a controlled lens on those questions and a way to measure progress rather than guess at it. We look forward to advancing this study, and to building the training recipes that turn these measurements into more reliable control. We're advancing this work now, and will share two more updates in the coming days. ### Part 2: Stress-Testing State-of-the-Art Models In Part 2, we'll put GUI-Perturbed to work as a benchmark. We'll take three state-of-the-art CUA models that share a base checkpoint but differ in their post-training recipes, and ask a sharper question than aggregate accuracy allows: as the scene and the instruction shift along controlled axes, which grounding capabilities hold, and which quietly fall apart? The perturbation design will let us pinpoint exactly where, and reason about why. ### Part 3: Can Fine-Tuning Close the Gap? In Part 3, we'll ask whether the fragility can be trained away. We'll fine-tune on GUI-Perturbed data and trace how performance moves as we vary the recipe and scale the data up, to see whether more data is enough or whether closing the gap takes something different. ## References 1. T. Xue et al., "An Illusion of Progress? Assessing the Current State of Web Agents," Oct. 08, 2025, arXiv: arXiv:2504.01382\. doi: 10.48550/arXiv.2504.01382. 2. "Operator System Card." Accessed: Feb. 27, 2026\. \[Online\]. 3. T. Xie et al., "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments," May 30, 2024, arXiv: arXiv:2404.07972\. doi: 10.48550/arXiv.2404.07972. 4. T. Xie et al., "Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis," Oct. 24, 2025, arXiv: arXiv:2505.13227\. doi: 10.48550/arXiv.2505.13227. 5. K. Cheng et al., "SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents," Feb. 23, 2024, arXiv: arXiv:2401.10935\. doi: 10.48550/arXiv.2401.10935. 6. K. Li et al., "ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use," Apr. 04, 2025, arXiv: arXiv:2504.07981\. doi: 10.48550/arXiv.2504.07981. 7. B. Gou et al., "Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge," Jul. 03, 2025, arXiv: arXiv:2506.21506\. doi: 10.48550/arXiv.2506.21506. 8. J. Y. Koh et al., "VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks," Jun. 06, 2024, arXiv: arXiv:2401.13649\. doi: 10.48550/arXiv.2401.13649. 9. T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang, "World of Bits: An Open-Domain Platform for Web-Based Agents," in Proceedings of the 34th International Conference on Machine Learning, PMLR, Jul. 2017, pp. 3135–3144. 10. X. Deng et al., "Mind2Web: Towards a Generalist Agent for the Web," Dec. 09, 2023, arXiv: arXiv:2306.06070\. doi: 10.48550/arXiv.2306.06070. 11. J. Yang et al., "GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies," Jun. 17, 2025, arXiv: arXiv:2506.14477\. doi: 10.48550/arXiv.2506.14477. 12. H. H. Zhao, K. Yang, W. Yu, D. Gao, and M. Z. Shou, "WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point," Feb. 22, 2026, arXiv: arXiv:2502.08047\. doi: 10.48550/arXiv.2502.08047. 13. T. Chen et al., "RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation," Aug. 27, 2025, arXiv: arXiv:2506.18088\. doi: 10.48550/arXiv.2506.18088. 14. L. Weng, "Domain Randomization for Sim2Real Transfer." Accessed: Feb. 27, 2026\. \[Online\]. 15. "Domain Randomization With Replicator — Getting Started With Isaac Sim." Accessed: Mar. 02, 2026\. \[Online\]. 16. J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, "Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World," Mar. 20, 2017, arXiv: arXiv:1703.06907\. doi: 10.48550/arXiv.1703.06907. 17. K. Yu, N. Yu, H. Wang, R. Yang, and H. Zhang, "How do visual attributes influence web agents? A comprehensive evaluation of user interface design factors," Jan. 29, 2026, arXiv: arXiv:2601.21961\. doi: 10.48550/arXiv.2601.21961. ## Get Involved At Fig, we believe reliable computer use requires models that understand why a GUI element serves a particular function, not just where it appears on screen. GUI-Perturbed is part of our broader work on control intelligence. You can reach out to us at contact@metarch.ai. ## Citation ```latex @online{gui_perturbed_technical_report_2026, title = {GUI-Perturbed: A Domain Randomization Dataset for GUI Grounding}, author = {Yangyue Wang and Harshvardhan Sikka and Yash Mathur and Tony Zhou and Jinu Nyachhyon and Pranav Guruprasad}, year = {2026}, url = {www.fig.inc/blog/domain-randomization-for-computer-control/}, note = {Part 1: Dataset & methodology} } ``` ### The Scaling Laws Are Breaking URL: https://www.fig.inc/blog/the-scaling-laws-are-breaking/ Last updated: 2026-06-08T01:36:10.000Z # *Part 1 of a series examining the need and search for new scaling laws.* *This piece reflects ongoing research at Fig. We welcome discussion and collaboration as we work to formalize these observations.* --- OpenAI's Orion consumed 15-30× more training compute than GPT-4\. By the scaling laws that have guided the field since 2020, it should have been revolutionary. Instead, internal reports suggest the model shows inconsistent improvements across task categories—sometimes performing worse than its predecessor on certain tasks \[12, 13\]. This is a fundamental break in the pattern that has driven hundreds of billions in infrastructure investment. The Kaplan \[1\] and Hoffmann \[2\] scaling laws promised predictable returns: more compute + more parameters + more data = reliably better performance. GPT-3 to GPT-4 followed this pattern beautifully. GPT-4 to Orion didn't. We've hit the constraint that Ilya Sutskever highlighted at NeurIPS 2024: "We have reached peak data" \[3\]. The internet's usable text—approximately 300 trillion tokens—is largely exhausted. Epoch AI projects that high-quality text data will be fully consumed by 2028, though current development suggests this timeline may be optimistic \[4\]. ## Three Paths, Three Dead Ends The field's response reveals how deeply we've internalized scaling orthodoxy. Each major approach tries to preserve the old paradigm rather than question it. ### Synthetic Data: The Recursive Trap If we've consumed the internet, why not generate our own training data? Train GPT-5 on GPT-4's outputs, bootstrap intelligence from itself. The information-theoretic constraints are unforgiving. Recent experiments show \[5, 6\]: - Model collapse begins by generation 5 when synthetic data exceeds 20% of the training mixture - Performance degrades 20-30% by generation 5, becoming unusable by generation 10 - Even Meta's successful use of synthetic data in Llama 3 relied on it as augmentation, not replacement You cannot bootstrap intelligence from nothing. The recursive loop provides no new information. ### Multimodal Scaling: The Complexity Explosion YouTube offers 20 billion hours of video, growing by 500 hours per minute. Surely this solves our data problem? The harsh reality: - Video's high redundancy compresses to perhaps 10-15 trillion useful tokens - Processing requires 10-100× more compute per token than text - Architectural modifications for video understanding reduce transfer efficiency to core language tasks We're not scaling the same thing anymore—we're building something fundamentally different with fundamentally different economics. ### Test-Time Compute: The $3,000 Question OpenAI's o3 achieves remarkable results: 88% on ARC-AGI (versus GPT-4's 5%), competitive programming performance ranking 175th globally \[8, 9\]. The method is conceptually elegant—let models "think" longer by generating extensive reasoning chains \[7\]. The economics are brutal: - o3-low: $17.50 per million tokens - o3-high: Over $3,000 per complex task - Average enterprise query cost: $87 - Customer willingness to pay: $0.10-1.00 But there's a deeper problem. Recent studies reveal that extended reasoning often makes models *less* reliable \[10\]: - Accuracy peaks at 1,000-2,000 reasoning tokens, then declines - Models confabulate increasingly elaborate incorrect justifications - They rationalize their way into errors they wouldn't make with quick responses We're not building intelligence—we're building very expensive overthinking machines. ## Why Traditional Scaling Had to Break The failure isn't accidental. It's structural. Traditional scaling laws rest on assumptions that no longer hold: **Static Deployment Assumption**: The laws assume models are trained, then frozen. But modern systems operate continuously, accumulating thousands of hours of interaction. A customer service agent after 10,000 conversations knows things that weren't in its training data. The framework cannot account for this. **Isolated Intelligence Assumption**: Scaling laws treat each model as self-contained. But deployed systems use tools, reference databases, call APIs. A 7B parameter model with browser access outperforms a 70B model without it. Traditional scaling is silent on this inversion. **Uniform Computation Assumption**: The laws assume every token gets equal processing. But intelligent systems should allocate compute based on difficulty—think harder about hard problems, less about easy ones. o3's uniform test-time scaling is like running a marathon at sprint pace. **Single Task Assumption**: Traditional scaling measures perplexity on held-out text. But we're asking models to browse the web, write code, use tools, and plan multi-step actions. The evaluation framework fundamentally misaligns with the deployment reality \[18, 19, 20\]. ## Toward Interaction Laws What we're discovering is that intelligence in real-world systems isn't a function of three variables (compute, parameters, data). It emerges from the interaction of many dimensions: - **Computational capacity** (traditional scaling) - **Experience accumulation** (learning through deployment) - **Tool augmentation** (compositional capabilities) - **Environmental complexity** (diversity of contexts) - **Temporal coherence** (persistence and memory) - **Network effects** (multi-agent dynamics) Each dimension has its own scaling properties. More importantly, they interact in ways we're only beginning to understand: - Small models with tools outperform large models without them - Experience accumulation shows phase transitions rather than smooth scaling - Multi-agent systems exhibit emergent capabilities unpredictable from individual agents We are tentatively calling these relationships **interaction laws**—principles that describe how different dimensions of capability interact and potentially amplify each other along axes of scale. ## Early Observations While we're still developing the formal framework, and will have more to share soon. To give you an idea of some patterns that are becoming clear: **Non-monotonicity**: Unlike traditional scaling's smooth curves, we observe discontinuous jumps. Systems seem to cross capability thresholds at specific experience levels. **Multiplicative Effects**: Tool integration doesn't add to capability—it multiplies it. The gain from tools scales with base model capability up to a saturation point we're still identifying. **Experience Efficiency**: Not all experience is equal. Diverse, challenging interactions provide more learning signal than repetitive tasks. This suggests curriculum design may be as important as scale. ## What's Working Now Three approaches are showing promise by implicitly acknowledging these multi-dimensional dynamics: **Adaptive Architectures**: Google's Gemini 2.5 selectively invokes expensive reasoning only when needed \[14\]. Most queries get fast, cheap responses; complex problems trigger deeper processing. This implicitly recognizes that uniform computation is wasteful. **Domain Specialization with Tools**: Anthropic's Claude Sonnet 4.5 achieves 70.3% on SWE-bench by deeply integrating with development environments \[15, 17\]. Rather than scaling the model, they're scaling the system's capabilities through tool mastery. **Efficiency Through Architecture**: DeepSeek R1 matches o1's performance at 3% of the cost using mixture-of-experts (671B parameters, 37B active) \[16\]. This suggests that how we scale matters as much as how much we scale. ## The Path Forward Understanding interaction laws will require rethinking our entire approach to AI development: **New Metrics**: We need benchmarks that capture multi-dimensional progress. How do we measure a system that learns from experience? That leverages tools? That persists across sessions? **New Architectures**: Models designed for experience accumulation and tool use from the ground up, not as afterthoughts. **New Training Paradigms**: Instead of maximizing compute on static datasets, we need to optimize across multiple dimensions simultaneously. We're at the beginning of this exploration. Over the coming months, we'll be sharing our research into these interaction dynamics—how to measure them, how to optimize for them, and what they mean for the future of AI. The breakdown of traditional scaling isn't a ceiling—it's a doorway. On the other side lies a richer, more nuanced understanding of intelligence that emerges not from brute force, but from the complex interaction of multiple capabilities. ## Citation Please cite this work as: ``` Sikka, H., & Fig AI Team. (2025, October 30). The scaling laws are breaking. Fig AI: Perspectives on Intelligence. https://blog.fig.inc/the-scaling-laws-are-breaking/ ``` Or use the BibTeX citation: ``` @article{sikka2025scaling, title={The Scaling Laws Are Breaking}, author={Sikka, Harshvardhan and {Fig AI Team}}, journal={Fig AI: Perspectives on Intelligence}, year={2025}, month={October}, day={30}, url={https://blog.fig.inc/the-scaling-laws-are-breaking/}, note={First in a series on interaction laws: why traditional scaling is failing and the need for multi-dimensional frameworks in AI development}, keywords={scaling laws, large language models, interaction laws, test-time compute, synthetic data, multimodal scaling, AI evaluation} } ``` ## References 1. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). *Scaling Laws for Neural Language Models*. arXiv preprint arXiv:2001.08361. 2. Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. V. D., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., ... Sifre, L. (2022). *Training Compute-Optimal Large Language Models*. arXiv preprint arXiv:2203.15556. 3. Sutskever, I. (2024, December). *Peak Data: Implications for Future AI Development* \[Keynote\]. Conference on Neural Information Processing Systems (NeurIPS) 2024, Vancouver, Canada. 4. Villalobos, P., Sevilla, J., Heim, L., Besiroglu, T., Hobbhahn, M., & Ho, A. (2024). *Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning*. Epoch AI. https://epochai.org/blog/will-we-run-out-of-ml-data 5. Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2023). *The Curse of Recursion: Training on Generated Data Makes Models Forget*. arXiv preprint arXiv:2305.17493. 6. Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Ahmed, H. R., Babaei, M., Baraniuk, R., & Lin, T. (2023). *Self-Consuming Generative Models Go MAD*. arXiv preprint arXiv:2307.01850. 7. OpenAI. (2024, December 20). *Deliberative Alignment: Reasoning Enables Safer Language Models*. https://openai.com/index/deliberative-alignment/ 8. OpenAI. (2025, January 17). *OpenAI o3 and o3-mini*. https://openai.com/index/openai-o3-and-o3-mini/ 9. ARC Prize Foundation. (2024, December 20). *OpenAI o3 Breakthrough*. https://arcprize.org/blog/openai-o3-breakthrough 10. Anthropic. (2025, July). *Inverse Scaling Properties of Chain-of-Thought Reasoning*. Anthropic Research Blog. 11. Hinton, G., Sutskever, I., Bengio, Y., LeCun, Y., Ng, A., Schmidhuber, J., & Russell, S. (2025, July). *Joint Statement on AI Reasoning Transparency*. Future of Humanity Institute. 12. Reuters. (2024, November 15). *Exclusive: OpenAI's next flagship model might not represent as big a leap forward as its predecessors*. Reuters Technology. 13. The Information. (2024, November 20). *OpenAI Shifts Strategy as GPT-5 Faces Delays*. The Information. 14. Google DeepMind. (2024, December). *Gemini 2.0: Our next era of models*. https://deepmind.google/gemini/ 15. Anthropic. (2024, October). *Introducing Claude 3.5 Sonnet and Claude 3.5 Haiku*. https://www.anthropic.com/news/3-5-models-and-haiku 16. DeepSeek. (2025, January 20). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv preprint arXiv:2501.12948. 17. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?*. arXiv preprint arXiv:2310.06770. 18. OSWorld Team. (2025). *OSWorld v3: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments*. In Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Datasets and Benchmarks Track. 19. H2O.ai. (2025, September). *GAIA Leaderboard Update: Progress and Persistent Challenges*. https://h2o.ai/blog/gaia-benchmark-2025/ 20. Xie, T., Zhao, Y., Li, Y., Wang, J., & Chen, L. (2025). *ARC-AGI-2: A Harder Benchmark for General Intelligence*. ARC Prize Foundation Technical Report. ### To Solve the Benchmark Crisis, Evals Must Think URL: https://www.fig.inc/blog/to-solve-the-benchmark-crisis-evals-must-think/ Last updated: 2025-10-26T14:18:51.000Z GPT-4 scored 95% on HumanEval. So did Claude. So did Gemini. But your production deployment still breaks on basic customer queries. We've collectively entered the what is fast becoming a dangerous phase of AI development: when benchmarks tell us nothing about what actually matters. Models have memorized the test set, RL has learned to game the metrics, and the gap between eval performance and real-world reliability has never been wider. The solution isn't harder benchmarks. It's a focused evolution for what we consider "evaluations". ## The Convergence of Frontier Models Look at any frontier model benchmark today. MMLU: everyone's above 90%. HumanEval: 95% across the board. GSM8K: solved. Even the "hard" benchmarks like GPQA are falling one by one. This should be good news. It's not. When every model scores identically, benchmarks stop being benchmarks. They become participation trophies. Worse, they create a dangerous illusion of capability. Your model aces MMLU's computer science questions but can't help a developer debug a race condition in production. It crushes mathematical reasoning benchmarks but fails to calculate your customer's multi-tier discount correctly. The contamination problem compounds this failure. Recent research from Sun et al. (2025) demonstrates that models like Aquila2 and Qwen reproduce verbatim training examples from MATH and GSM8k¹. GPT models perform significantly better on coding problems released before their training cutoff². When LLMs barely improve over majority baselines on truly uncontaminated tasks³, what's measured is no longer capability, but something akin to memorization. RL saturation is also difficult. Give any competent RL system enough iterations on a fixed benchmark and it will find a way to maximize the score. Not by becoming more capable, but by exploiting the specific patterns that benchmark rewards. This is Goodhart's Law4 at scale: when a measure becomes a target, it ceases to be a good measure. A practical anecdote: A team we recently worked with through [Fig Labs](https://fig.inc/?ref=fig.inc) evaluated a model scoring \~97% on coding benchmarks that completely failed to refactor a relatively simple front end component (written in React). Why? The benchmarks test algorithmic puzzles. Real codebases need understanding of dependencies, side effects, and business logic. The model had learned to solve puzzles, not write software. ## Static Benchmarks have Challenges Static benchmarks made sense in a certain time, sometime prior to 2023—when models couldn't learn from experience, couldn't access external tools, and certainly couldn't be trained specifically to beat evaluations. The fundamental assumption was that evaluation was measurement, not competition. Today's models undergo continuous refinement through RLHF, fine-tuning, and increasingly, online learning. They don't just take tests—they study for them. And when you're testing something that can study specifically for your test, static evaluation becomes theater. Academic benchmarks compound the problem by measuring yesterday's challenges with yesterday's assumptions. Hugging Face acknowledged this directly when launching their Open LLM Leaderboard v2: "Models began to reach baseline human performance on benchmarks like HellaSwag, MMLU, and ARC, reducing their effectiveness in distinguishing model capabilities"5. The disconnect is stark. Labs celebrate benchmark improvements while enterprises struggle with basic reliability. It's like claiming your car is race-ready because it aces emissions tests, while it stalls every time you need to merge onto the highway. ## Evaluations as Intelligent Systems The paradigm shift is simple but profound: evaluations must become intelligent systems themselves. Instead of fixed test sets, we need evaluators powered by sophisticated generation models that create infinite scenarios. These aren't random perturbations—they're adversarial environments specifically crafted to probe model boundaries. The evaluator learns what makes models fail and systematically explores that space. Consider LiveBench, introduced by White et al. (2024)6. It releases new questions monthly, drawn from recent arXiv papers, news articles, and datasets—all postdating model training cutoffs. Questions are scored automatically against objective ground-truth values. The result? Top models achieve below 70% accuracy, maintaining discriminative power even as capabilities improve. This is just the beginning. The evolution happens at multiple levels: 1. **Procedural generation** ensures every model faces unique challenges. MCPEval demonstrates this by automatically synthesizing tasks from available tool APIs, then deploying "frontier agents" to verify executability7. Like roguelike games where every run is different, evaluations become inherently resistant to memorization. 2. **Dynamic difficulty** maintains discriminative power as capabilities improve. When models start succeeding, the evaluator increases complexity. The benchmark evolves with the frontier, always challenging, never saturated. 3. **Adversarial discovery** actively searches for weaknesses. Recent work from Anthropic shows models generating attacks against themselves in loops, with each iteration uncovering novel failure modes8. It's the difference between random quality control and having a dedicated red team that never sleeps. ## The New Evaluation Stack Three pillars will define the next generation of evaluation infrastructure: ### Living Benchmarks Benchmarks that update continuously from real-world sources. Questions drawn from this week's research papers, not datasets frozen in 2021\. Code challenges from actual GitHub issues, not toy problems. Customer service scenarios from production logs, not synthetic dialogues. LiveBench provides the template: temporal isolation through post-training-cutoff data, automated objective scoring, and monthly updates6. But this is version 1.0\. The next generation will incorporate production telemetry directly, turning every deployment into an evaluation opportunity. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2025/10/Screenshot-2025-10-26-at-7.09.41---AM.png) [LiveBench.ai](https://livebench.ai/?ref=fig.inc#/) ### Learned Simulators World models that generate evaluation environments, not just test cases. These simulators understand the task space deeply enough to create meaningful variations. They don't just permute inputs—they understand what makes a task hard and systematically explore that difficulty space. Imagine evaluating a code generation model. A learned simulator doesn't just vary function names. It generates scenarios requiring genuine architectural decisions: handling race conditions, managing state across microservices, refactoring with backward compatibility constraints. Each test is novel, yet grounded in real engineering challenges. The adversarial dynamic creates an evolutionary arms race. Models improve to beat evaluators. Evaluators evolve to find new weaknesses. Recent research on Deep Adversarial Automated Red Teaming (DART) shows this approach reducing violation risks by 53.4%9—improvements impossible with static evaluation. ### Wild Deployment as Ground Truth The ultimate evaluation is production. Real users, real tasks, real consequences. No synthetic benchmark matches the complexity and unpredictability of actual deployment. Discord's deployment of their AI capabilities provides a masterclass in production-driven evaluation10. They implemented continuous telemetry, passive moderation to detect adversarial trends, and quantitative risk measurement across stakeholders. Every interaction became data. Every failure became a future test case. ![](https://storage.ghost.io/c/a0/84/a0848f29-ea39-4b8c-a26d-c283a5a64f0c/content/images/2025/10/image.png) [Diagrammatic Overview of Discord's 2 Phase AI-assisted evaluation](https://discord.com/blog/developing-rapidly-with-generative-ai?ref=fig.inc) This closes the loop completely. Models train on human feedback, deploy to production, generate evaluation signal, which feeds back into both model improvement and evaluator evolution. The entire system learns from reality, not approximations. ## The Economic Transformation Here's what most people miss: evaluations aren't just testing infrastructure. They're economic infrastructure. They literally define what AI can do in the economy. Brandon from Mercor captures this perfectly: "Evals are the new PRD." They specify capabilities as precisely as any product requirements document. If you can't evaluate it, you can't deploy it. If you can't deploy it, it has no economic value. This reframes the entire AI development stack. The bottleneck isn't model capability—it's evaluation coverage. We have models that could automate vast swaths of knowledge work today, but we can't deploy them because we can't reliably evaluate whether they'll work. Consider customer support automation. The model capability exists. What's missing? Evaluations that cover the full range of support scenarios, edge cases, and failure modes. Without comprehensive evaluation, deployment is gambling with your brand. Living evaluations transform this dynamic. Recurring workflows become one-time evaluation setup costs. Instead of manually reviewing every model output, you build an evaluator once and run it forever. The variable cost becomes fixed. The uninsurable becomes predictable. This is how AI reaches the entire economy. Not through better models alone—through better evaluations that make deployment safe, predictable, and economically viable. ## A Reality Check for Implementations Building dynamic evaluation systems isn't free. Our early experiments (more on this soon!) suggest several technical constraints: **Computational overhead**: Procedural generation and adversarial search increase evaluation costs by 10-100x compared to static benchmarks. But this is still cheaper than production failures. **Infrastructure complexity**: Version control for evaluation environments, contamination detection, automated validation pipelines—the engineering lift is substantial. LiveBench's infrastructure provides a blueprint, but implementation remains non-trivial⁶. **Standardization tension**: Dynamic evaluation inherently resists standardization. This could fragment the field, making cross-model comparisons difficult. We need new frameworks for comparing models evaluated on different (but theoretically equivalent) test distributions. ## What This Means for AI Development There are implications that cascade through the entire AI stack: **No more benchmark overfitting.** When evaluations evolve continuously, gaming becomes impossible. Models must develop genuine capabilities, not benchmark-specific tricks. **Continuous evaluation pipelines.** Evaluation isn't a gate before deployment—it's integrated into deployment itself. Every production interaction generates evaluation signal. **Real capability discovery.** Intelligent evaluators don't just test known capabilities—they discover unknown ones. They find emergent behaviors, unexpected strengths, and hidden failure modes. The competitive dynamics shift fundamentally. Organizations that win won't necessarily be those with the biggest models or most compute. They'll be those with the most sophisticated evaluation infrastructure. The ability to rapidly and reliably assess capabilities becomes the core differentiator, and we're seeing early signs of this already. The most successful AI deployments aren't using the most advanced models—they're using the most comprehensive evaluation systems. Prosus built their evaluation using millions of real queries from their Toqan assistant11. They can deploy with confidence because they know exactly how their models will perform. ## The Path Forward The timeline is aggressive but achievable. By 2026, static benchmarks will be obsolete for frontier development. Organizations still relying on them will be unable to compete—not because their models are worse, but because they can't prove their models are better. The technical pieces exist. Learned Environmental modeling (colloquially called "World Models") is on a promising path. Procedural generation works. Production telemetry systems work. What's needed is integration—bringing these components together into unified evaluation infrastructure. The economic incentive is massive. Companies that crack evaluation unlock the entire AI economy. They can deploy where others can't, automate what others won't, and scale while others stall. But this isn't just about individual companies. When evaluations evolve, AI development accelerates. When we can reliably assess capabilities, we can safely deploy them. When deployment generates evaluation signal, the whole system improves. The next breakthrough in AI won't necessarily come from a larger model or better training technique. It will come from evaluations that think, adapt, and learn. Evaluations that discover what we didn't know to test for. Evaluations that evolve faster than models can overfit. The benchmark crisis isn't just a technical problem—it's the key that unlocks AI's economic potential. Static benchmarks are dead. The question isn't whether your model scores 95%. It's whether your evaluator is smart enough to find the 5% that matters. When that happens, when evaluations truly evolve, everything changes. Not because models get better, but because we can finally trust them enough to use them. *Fig AI is actively developing tools for dynamic evaluation as part of the full stack for next generation AI systems. We welcome collaboration from teams working on similar challenges. You can learn more about our work at fig.inc, and reach out to us directly* [*here.*](mailto:contact@metarch.ai) --- ## References 1. Sun et al. (2025). "Benchmark Data Contamination of Large Language Models: A Comprehensive Analysis." *arXiv preprint*. 2. Ravaut et al. (2024). "The Evolving Landscape of LLM Evaluation: Contamination and Best Practices." *ACL*. 3. Roberts et al. (2024). "Test Set Contamination in Large Language Models: Evidence and Implications." *ICML*. 4. Goodhart, C.A.E. (1975). "Problems of Monetary Management: The U.K. Experience." Papers in Monetary Economics. Reserve Bank of Australia. Later popularized as "Goodhart's Law" by Strathern, M. (1997) in "Improving ratings: audit in the British University system." 5. Hugging Face Team (2024). "Open LLM Leaderboard v2: Addressing Benchmark Saturation." *Technical Report*. 6. White, C., et al. (2024). "LiveBench: A Challenging, Contamination-Limited LLM Benchmark." *NeurIPS*. 7. Liu et al. (2025). "MCPEval: Dynamic Task Generation for LLM Agent Evaluation." *arXiv preprint*. 8. Anthropic (2024). "Automated Red Teaming: Models Testing Models." *Technical Report*. 9. Jiang et al. (2024). "DART: Deep Adversarial Automated Red Teaming." *ICML*. 10. Discord Engineering (2024). "Deploying AI at Scale: Lessons from Production." *Internal Case Study*. 11. Prosus AI Team (2024). "ProLLM: Real-World Benchmarks from Million-Scale Deployments." *Technical Blog*. ## Citation Please cite this work as: ``` Sikka, H., & Fig AI Team. (2025, October 26). To solve the benchmark crisis, evals must think: Dynamic evaluation for foundation models. Fig AI: Perspectives on Intelligence. https://blog.fig.inc/to-solve-the-benchmark-crisis-evals-must-think/ ``` Or use the BibTeX citation: ``` @article{sikka2025evals, title={To Solve the Benchmark Crisis, Evals Must Think: Dynamic Evaluation for Foundation Models}, author={Sikka, Harshvardhan and {Fig AI Team}}, journal={Fig AI: Perspectives on Intelligence}, year={2025}, month={October}, day={26}, url={https://blog.fig.inc/to-solve-the-benchmark-crisis-evals-must-think/}, note={Perspective on addressing model capability saturation through dynamic, adversarial, and production-driven evaluation systems}, keywords={large language models, evaluation, benchmarks, contamination, dynamic evaluation, adversarial testing, foundation models} } ```