Did your agent get worse last week?

Trajecta is regression testing for browser agents. Define a task suite, run your agent against it, and diff the trajectories across runs and model versions. You will know the moment a model upgrade silently breaks your agent.

How it works

Three steps. No database, no hosted service to babysit. Runs locally today.

01

Define tasks

A plain YAML file: starting URL, goal in words, and assertions like URL reached, text present, or extracted data matching a schema.

02

Run the suite

One CLI command runs every task and records the full trajectory: actions, DOM snapshots, tool calls, tokens, latency, errors.

03

Diff the results

Compare any two runs. New failures, fixed tasks, changed action sequences, cost and latency deltas. Exits non-zero on regressions, so it drops straight into CI.

What a regression looks like

Baseline: 3/3 passing. Candidate model: one task silently changed behavior.

# tasks.yaml
- name: extract-gdp-table
  startUrl: https://en.wikipedia.org/...
  goal: Extract the top 5 countries by GDP
  assertions:
    - type: schema
      # extracted rows match expected shape
- name: pricing-page
  startUrl: https://github.com/pricing
  goal: Reach the pricing page
  assertions:
    - type: url_contains
      value: /pricing
$ trajecta diff --base v1 --candidate v2

PASS  extract-gdp-table
PASS  pricing-page
FAIL  checkout-flow        new failure

action sequence changed: 14 -> 22 steps
tokens: 8.1k -> 12.4k (+53%)
latency: 41s -> 68s

exit 1: 1 new failure

The candidate model took a different, longer path and failed a task that used to pass. Without the diff, you would have shipped it.

Built for teams shipping browser agents

If any of these sound familiar, Trajecta is for you.

Sample output

Every diff also saves a machine-readable artifact at runs/diffs/<base>--vs--<candidate>.json, so CI or your own tooling can act on it.

{
  "base": "baseline",
  "candidate": "candidate",
  "newFailures": ["github-pricing-plans"],
  "fixed": [],
  "changedSequences": [
    {
      "task": "wikipedia-gdp-top5",
      "base": ["goto", "observe", "extract"],
      "candidate": ["goto", "scroll", "observe", "extract"]
    }
  ],
  "costDeltas": [
    {
      "task": "wikipedia-gdp-top5",
      "tokensDelta": 150,
      "latencyDeltaMs": 750
    }
  ],
  "totalsDelta": { "tokens": 120, "latencyMs": 1700 }
}

FAQ

Short answers to the questions everyone asks first.

What it does

What does Trajecta actually do?

It records everything your browser agent does on a task suite, then diffs trajectories across runs and model versions. When a new model takes a worse path or fails a task that used to pass, you see it as a clear diff instead of finding out in production.

Why not manual

How is this different from spot-checking runs myself?

Spot checks don't scale and don't catch slow drift. Trajecta replays the whole suite on every change and exits non-zero on regressions, so it plugs into CI and answers "did the agent get worse this week?" without anyone watching runs by hand.

Pricing

How much does it cost?

Trajecta is in early access and free to try right now. The local CLI runs on your machine with no hosted service. Longer term pricing is still being figured out, early users will have a say in it.

Fit

Who is it for?

Teams shipping browser agents who change models, prompts, or tools and want to know what broke before users do. If you have ever upgraded a model and quietly shipped worse behavior, that is the pain this fixes.

Stop guessing. Start diffing.

Trajecta is in early access. Tell us what your agents do and we will get you running.

Get early access

Built by someone who lives this

Sai Preetham Reddy Leburu

SDE at Amazon, building MCP servers and evaluation frameworks for production agents. Trajecta is the tool I wished existed every time a model upgrade quietly changed what my agents did.

LinkedIn  ·  GitHub  ·  Trajecta source code  ·  hello@gettrajecta.com