Visual Intelligence & Computer Use

Computer use to
see and act

Proteus is an AI engineering lab building the next generation of computer use agents, combining fast visual understanding, efficient context compression, and RL to let machines learn to operate multi-interface workflows on any device, autonomously.

Reasoning has
outpaced vision

Coding models and reasoning models are far beyond most humans today, but that is not true for computer use. We think this is due to slow image encodings, bloated context windows, wrong actions removable only via custom RL, and no ability to act continuously over long horizons. We want to fix that.

01 — Fast perception

Small models for rapid action

Fine-tuned compact models for instant visual grounding and action prediction, paired with large VLMs that handle high-level planning and reasoning at different latencies simultaneously (think: fast model acts and collects data to generate rollouts, while slow model determines correct actions).

02 — Efficient encoding

Efficient representations

We use fast OCR, segmentation, and video models to build optimized visual representations that eliminate the redundancy of feeding raw screenshots into frontier models -- leading to faster inference, longer horizons, and lower cost.

03 — Continuous operation

Long horizon planning

Agents should learn diverse, long-horizon workflows from video and be able to navigate real interfaces on desktop, mobile, and the web. Annotated DOMs or HTML should be only used when learning new workflows, not at inference time.

Test on any device

We built lightweight action interfaces for desktop, web, and mobile, with a shared brain to handle reasoning and reflection. Here's a preview.

Real-time Android agent · Prompt: "Find me an Airbnb in San Francisco for 10 people this weekend. Money is no object, pick the best place." · Real-time, clicks where needed, handles popups, recovers from failure, and asks the user for clarification when intent is ambiguous.

Autonomous mobile agents

Rapid computer use on Android and iPhone. The agent sees the screen, reasons about what to do, and takes action, navigating apps, filling forms, and completing multi-step tasks with no predefined scripts.

Powered by compact vision models for fast grounding and frontier LLMs for planning, with custom optimizations for efficient visual context that keep latency low and accuracy high.

Windows

Native computer use on Windows desktops. Operates Win32 and UWP apps, manages files, runs multi-step workflows across the OS.

Ubuntu

Full desktop automation on Ubuntu Linux. Navigates GNOME, terminal, and GUI apps with the same visual understanding pipeline.

macOS

Seamless computer use on macOS. Controls native apps, Finder, and system interfaces through real-time screen understanding.

Deep expertise in AI

Our team's published research prior to Proteus.

We also contributed to work on decoding visual imagery via fNIRS, adversarial examples, and reinforcement learning.

Our team comes from

Everything is computer

We partner with teams building computer use agents, visual AI infrastructure, and autonomous systems that need to see and act in the real world.