
Agent Interface
Give AI agents a better way to use computers
52 followers
Give AI agents a better way to use computers
52 followers
AI agents need better computer tools, not just better models. Agent Interface is an open-source layer built to cut repeated screenshots, model calls, and waiting. It reuses learned interactions, runs action-and-feedback loops locally, and asks the model when fresh judgment is needed. Start with the runnable desktop research preview, and follow the Astra/Freedoom experiments exploring control in a world that doesn't pause while AI thinks.







"The world doesn't pause while the AI thinks" is the honest framing most computer-use demos skip. Separating the reusable procedure from the current binding is the right cut. Question: when the local loop revokes an action because fresh state disagrees with the plan, what does the model get back, a diff of what changed or the whole new screen? Asking because that's the expensive part in every agent I've run.
@thicreator we actually tested this, and the answer ended up being “neither diff nor full screen as a fixed rule.”
The first goal is to avoid going back to the model at all. If the invalidation is locally decidable, the runtime repairs/revalidates from fresh current evidence and recommits locally. In one matched Chromium repair experiment, that path took about 110 ms, while model reacquisition took about 8.1 s and roughly doubled the input tokens.
If the change crosses a semantic boundary the local runtime cannot fully validate, then it yields back to the rich model. And importantly, even after the model returns, we don’t treat that answer as fresh authority — we take another passive current observation and revalidate the target before input.
We also tested what visual context to send back. On Chromium, current-only, prior full-frame, and an action-grounded crop were all correct; the crop saved ~11% versus full history, but it was slightly worse than just sending the current state. Then on OpenTTD, the crop failed one of two transitions while current-only and full context were both correct.
So the rule we’ve converged toward is: local repair when semantics are locally complete; otherwise yield, and give the model the smallest current evidence that is actually sufficient — not a fixed “always diff” or “always screenshot” policy.
@unjuno 110 ms local repair versus 8 seconds and double the tokens to go back to the model is the whole argument for the project in one line. Thanks for the detail, that's a better answer than most papers give.