Eval suite
Run #284 · on every docs sync
Forward reads your API documentation and ships a production-grade AI co-pilot you embed in your product. Your customers say what they need. The agent handles it through your own API.
Raphael, The School of Athens, 1511 · Public domain, Wikimedia Commons
Wide Renaissance fresco in the manner of Raphael's School of Athens: a grand vaulted marble hall filled with scholars in classical robes debating and studying, warm plaster and stone tones, soft diffuse light, gentle fade to white at the edges, museum-scan quality, no text, no watermark.
+ 210 more, indexed and searchable
sk_live_…f92a · production
Reads what you already have. OpenAPI or Swagger, docs sites, Postman, GraphQL, Markdown.
Your API stays in charge. Signed sessions, your permission checks, confirmation gates on writes.
Measured from day one. Traces, evals, and regression gates are built in, not bolted on.
Most SaaS products lose users in the first ten minutes, not because the product is bad, but because the user can't find the thing they came for.
→ Get in touchDrop in a URL to your API reference, OpenAPI spec, or docs site. Forward crawls it, resolves schemas, and maps every endpoint to something a user would actually ask for.
Tools get generated, named, and grouped. Auth is wired to your session context. Write actions get confirmation gates by default. You review the catalog before anything ships.
A single SDK component, four surfaces to choose from. Your app passes a signed session token, and your API stays the source of truth for who can do what.
J. M. W. Turner, The Red Rigi, 1842 · Public domain, Wikimedia Commons
Turner watercolor of a mountain across a still alpine lake at dawn: luminous layered washes, pale rose and gold light, soft blue mountain silhouette, atmospheric mist, delicate lifted highlights, aged paper texture, no text, no watermark.
+ 211 more endpoints, indexed and searchable
<ForwardCopilot mode="bubble" session={token}/>One component. Four surfaces.
Same agent, same session, same tool catalog. Pick the surface that fits how your users already work. it's one prop on the component.
John Constable, Cloud Study, 1822 · Public domain, Wikimedia Commons
English Romantic watercolor cloud study in the manner of Constable: soft grey-blue cumulus washes bleeding into warm white paper, loose brushwork, airy negative space, subtle paper grain, no text, no watermark.
Last 14 days · production
An agent you can't measure is a liability. Every conversation is traced, every tool call is logged, and your eval suite runs on every change.
Every turn, tool call, token and cost, replayable exactly as the user hit it.
Promote any real conversation into a test case in one click.
Every docs sync re-runs the suite and blocks rollouts that drop your scores.
Page on failed tasks and wrong tools, not CPU. Slack and webhooks included.
Run #284 · on every docs sync
conv_9f13c · 2.4s total
"invite alice as admin, no billing"0ms214 tools → 4 candidates → invite_member+180ms · 1.2k tokensPOST /v1/workspaces/ws_82f/members · 201+840msConfirmed to user · task marked resolved+2.4s · $0.004Context gets polluted, the wrong tool gets picked, and users stop trusting it. Forward is the engineering between "it demos" and "it works."
Where activation leaks
Paste a link to your API documentation and your email. We'll review it and send back a private co-pilot link before you write any integration code.
Free to try · No credit card · Your co-pilot link lands in your inbox