Early access
Your AI agent says it’s done. Prove it.
An independent verification layer for software built by AI coding agents. It tests behavior, surfaces hidden assumptions, and flags what you didn’t know to ask about.
Feature complete.
Removed member can still access project data.
- Claim
- Removed members cannot access project data.
- Check
- Removed member sends an authenticated request.
- Result
- 200 OK. Expected 403.
- Request
- GET /projects/42, session of a removed member
- Response
- 200 OK
- Expected
- 403 Forbidden
- Given
- A member who was removed from project 42.
- When
- They request GET /projects/42 with their old session.
- Expect
- 403 Forbidden.
- Task
- Fix AUTH-07: removed members can access project data.
- Status
- open
- Fix
- Applying fix…
- Re-run
- AUTH-07: removed member sends an authenticated request.
- Result
- 403 Forbidden. The invariant holds.
The problem
AI got very good at writing code. That created a new problem.
Coding agents build features in minutes. You’re still the one who has to find out whether those features keep their promises.
- Human writes code
- Human reviews code
- Human tests product
- AI writes code
- AI says “done”
- Human asks: “Is it actually done?”
The hard part isn’t syntax errors. It’s what the product quietly gets wrong, or was never asked to get right:
The core idea
You don’t have to know what you forgot to ask.
Your coding agent knows what it built. Assay checks what it promised, including the promises nobody wrote down. You shipped a team app. But did anyone verify:
01 · ACCESSCan a removed member still access a project?
02 · DATAIs the data actually gone after a delete and a refresh?
03 · PRIVACYDoes that “just a font” request send your visitors’ IP addresses to a third party?
04 · ANALYTICSWhat can your session replay record from a form field?
05 · EMAILIf you email your waitlist, are unsubscribe and sender requirements covered?
06 · ASSETSWhere did that image come from, and are you allowed to use it?
07 · DISCOVERABILITYDoes the site have the important discovery files it needs, such as robots.txt, sitemap.xml, or llms.txt?
How it works
Claim. Check. Evidence.
Discover claims
Reads the application and its stated behavior to find what it promises, explicitly and implicitly.
Turn claims into checks
Important claims become invariants. Assay generates adversarial scenarios designed to break them.
Run them, keep the evidence
Checks execute against the application where possible. Every result carries the request, the response and the outcome.
Keep what matters
Useful invariants persist and re-run on future changes. What can’t be verified automatically is surfaced for human review.
Why it’s different
Here’s the claim. Here’s the test. Here’s what happened.
Assay doesn’t tell you your code looks suspicious. It states a claim, tries to break it, and shows you the result. You decide when to ship. Then act on the finding, fix it, and verify the fix.
- Agent said
- “Implemented project permissions.”
- Invariant
- A removed member can never read project data.
- Generate
- Adversarial checks designed to break the claim.
- Run
- GET /projects/42 with the old session of a removed member, against the application.
- Evidence
- removed member -> authenticated request -> 200 OK
- Status
- INVARIANT VIOLATED
Example output
Evidence, not vibes.
One run, many kinds of promises: behavior, privacy, compliance signals, site configuration, and assets. Privacy and compliance findings are review flags, never legal conclusions.
response: 200 OK (expected 403 Forbidden)
sent: visitor IP address and user agent. Review flag, not a legal conclusion.
Interactive demo · simulated results. Assay is in development; this is the report format we are building toward.
Differentiation
Not another coding agent.
Your coding agent is optimized to build the feature you asked for. Assay is optimized to challenge the result. Each layer has a job.
Assay verifies itself
This page was checked first.
If we say software should be verified, our own should survive it. Only what we verified is listed here, plus the one issue we found. The checks are in the site repository (npm run check).
Early access