the gateprivate beta

The AI analyst that has to prove it's right.

Evalyst connects to your databases read-only — and writes its own connectors for everything else, from a description in plain language. Ask it what you'd otherwise dig out by hand; every answer is checked against ones your team confirmed before it reaches a dashboard.

Access is granted, not instant — confirm your email and we'll come back to you. One click, no password, link good for 15 minutes.

read-only · your rows stay in your database · not used to train models

the output

Answers become dashboards — pinned to the data.

illustrative data
gross margin % · daily
40 50 60 70 80 90 Jul 01Jul 08Jul 15Jul 22Jul 29 avg 69.7%
spend usd · by service
$0 $400 $800 $1.2k Jul 05Jul 14Jul 23
compute storage egress
open balance · by account
billing@moon-mining.test
$18.4k
ap@robot-bakery.test
$13.1k
finance@time-travel.test
$8.6k
ops@cloud-castles.test
$4.2k
ar@unicorn-rental.test
$2.5k
admin@kraken-cafe.test
$1.7k
billing@yeti-logistics.test
$1.2k
others
$3.3k
collectible stale <1y written off

refreshes re-run the saved artifact — a parameterized query, a cross-source workflow, or a rollup. no model in the loop, so these numbers move only when the data moves.

no llm at refresh

Your dashboards don't call a model.

The agent writes the artifact once — a parameterized query, a workflow that reduces at every source it touches and joins the small results, or a materialized rollup. After that it's code, not a prompt.

Everything recurring runs the same way: widget refreshes, scheduled connector pulls that append history for day-over-day deltas, and alert rules that evaluate a condition and notify you. None of them call a model.

The same question gives the same answer

A dashboard that re-asks a model every morning can quietly change its mind about last month's revenue. This one can't: the number moves only when the data moves.

Viewers cost nothing to add

Nobody's read of a chart triggers inference, so there's no reason to charge by the seat. Invite the whole company.

bring your own database

Connect what you have, including the awkward parts.

Four engines are native, reached directly or through your bastion. Anything that isn't a database gets a connector the agent writes for you — that's the next section.

  • postgres native driver
  • mysql native driver
  • snowflake native driver + pinned warehouse
  • bigquery native driver + scan cap
  • ssh bastion tunnel, key scoped to one database
  • http agent-written connector, sandboxed

Row caps and statement timeouts everywhere, a bytes-billed ceiling on BigQuery that fails an over-budget query before it scans, and a pinned warehouse on Snowflake.

the long tail

No connector for it? Describe it in chat.

Native drivers cover the databases. For everything else — the billing API, the internal service, the vendor with one undocumented endpoint and no client library — you say what you need in plain language and the agent writes the connector for it.

Then it has to prove it works. The code it writes runs sealed in a container of its own, able to reach nothing it wasn't handed and with no way onto the network unless it asked for one. It goes live only if it really connects and really comes back with fields — and if it doesn't, you get the error instead of a connector that quietly half-works.

  • ✓The model never sees your key. It declares the field name; you type the value into a form that encrypts it.
  • ✓When the endpoint changes underneath you, a monitor catches the drift and repair rewrites the connector — through the same gate, so nothing is reactivated for looking fixed.
  • ✓What it pulls lands as a table you can join against your databases in a single query.
new connectorillustrative
you

connect to our billing API at api.acme.test — I need invoices and their status. it takes an API key.

evalyst

Wrote billing-invoices — three methods. Validating in the sandbox.

  • ✓test() → ok
  • ✓discover_schema() → 14 fields
  • •needs BILLING_API_KEY — you enter it; I never see the value

active ready to ingest, and joinable in a workflow

api · mcp

Ask your data from the tools you already use.

Everything the chat here can do — ask a question, pin a chart, set a schedule running — is reachable over a public API and an MCP server. Point your own client at your workspace and your data room answers there instead.

Where your client can draw them, charts come back as charts rather than a wall of numbers.

The same rules travel with it. Read-only stays read-only, the semantic model is still what the answers mean, and the gate still stands between a changed metric and a number anyone reads. A different window, not a different set of guarantees.

  • Claude Desktop
  • Claude Code
  • ChatGPT
  • any MCP client
mcp · your clientillustrative
{
  "mcpServers": {
    "evalyst": {
      "url": "https://mcp.evalyst.ai",
      "token": "<your workspace token>"
    }
  }
}
you, in your own client

what did gross margin do last month?

evalyst
40 50 60 70 80 90 Jul 01Jul 08Jul 15Jul 22Jul 29 avg 69.7%

The same chart the data room draws — rendered by your client, not pasted as text.

why this is the hard part

The dangerous answer isn't “I don't know.” It's “$0.”

A wrong number doesn't error. It renders — and numbers are believed. It gets pasted into a board deck and argued from for weeks, because nothing about it looks broken.

Nobody catches a silent error by reviewing SQL, because the SQL looks right. You catch it by having a number someone already signed off on — and checking against it every single time.

One silent error in a run of otherwise perfect answers. It's the only one of the four states that never announces itself.
how one happens

A column called free marks whether a stream was billable. Every row written before 2023 is NULL — the flag was added later, nobody backfilled it.

-- looks right. drops every pre-2023 row.
SELECT SUM(amount) FROM streams
WHERE free = 0 AND year = 2025;

→ $0 · query ran · nothing errored · chart rendered

NULL = 0 is never true, so an entire year of revenue silently drops out. That's the failure Evalyst is built around.

attach · draft · gate

Three steps, each one a real gate.

  1. 01

    Attach read-only — and we check that you did

    Postgres, MySQL, Snowflake or BigQuery. A source whose credentials can write does not activate, and the error tells you exactly which grant to fix. For SSH we generate the tunnel key ourselves, scoped to that one database.

  2. 02

    The agent drafts what your data means — you correct it

    It reads your schema and checks each hypothesis with a real query: which column is money, which dates are authoritative, which flags are NULL-able. Then it hands you the draft. You fix what it got wrong — every save is a new version.

  3. 03

    Your confirmed answers become the gate

    The agent proposes questions with real numbers attached; you confirm, correct or reject. From then on that bank stands between every change and your dashboards: one silent error, and the change doesn't ship.

measured, not asserted

Every source carries its own scorecard.

Accuracy, silent-error count, how many questions are in the bank, when it last ran. It sits on the source and travels with every answer that source produces. When someone asks whether they can trust the number — this is what you show them.

Not a claim. A measurement.

prod-postgres read-only verified
sample layout
ACCURACY —
SILENT ERR —
BANK SIZE —
LAST RUN —

figures populate from real eval runs against your source — never asserted, never mocked.

you're handing us production credentials

So here's exactly what we do with them.

  • ✓Read-only is proven when you attach the source, not assumed at query time.
  • ✓The agent has no shell. Its only tools are the ones we expose — no commands, no files, no arbitrary network.
  • ✓Credentials go from an admin-only form into an encrypted vault. The model sees field names, never values.
  • ✓Every query is parsed and rejected unless it's a single read-only SELECT in your engine's dialect.
  • ✓One workspace's data is isolated from another's at the query layer and again in the database itself.

How it's built →

who this is for

Three kinds of team, one shared problem.

Different stacks, different headcounts — and in all three, the meaning of a number lives in somebody's head, where nothing can check it.

  • usage-billing saas

    Real revenue data, no data team

    Your numbers live in Postgres or MySQL, half the revenue story lives in a billing API nobody wrote a connector for, and the person answering “how much did we make on X last quarter” is a founder with a SQL client open at midnight.

  • finance & ops

    The close that means joining three systems by hand

    Every month somebody reconciles the production database against the payment processor and a billing API, in a spreadsheet, the same way as last month. That join becomes a workflow that runs itself — and reports the same number twice.

  • teams with a warehouse

    Pipelines you trust, definitions you don't

    You already have the warehouse, the models and the dashboards. What you don't have is one agreed answer to what counts as an active customer — so three dashboards say three things, and the argument is about whose query was right rather than what happened.

objections

The questions people actually ask.

Can it write to my database?

No. Evalyst refuses to activate a source whose credentials can write — it checks the engine's own view of effective privilege and runs a probe that writes nothing but makes the server reveal what it would allow. On top of that, every query is parsed and has to be a single read-only SELECT, and the agent has no shell: its only tools are the ones we expose, and running commands is not among them.

Can it connect to our internal thing?

If you can describe it, probably. Databases get native drivers; anything with an API gets a connector the agent writes from your description, tests in a locked-down sandbox, and registers only if it really connects and really returns fields. If it fails, you get the error — not a connector that quietly half-works.

Is my data used to train a model?

No. Queries run through Anthropic's API under commercial terms that exclude training on your inputs and outputs. Your rows stay in your database — we hold query results long enough to answer and to snapshot the dashboards you pin.

What happens when it gets something wrong?

That is the part we built the product around. Every answer lands in one of four states: correct, abstained, wrong, or a silent error — confident and wrong. The last one is the only genuinely dangerous state, so changes to what your metrics mean are blocked unless the bank of questions your team confirmed still passes with zero of them.

Who can see the data?

People you invite, under a role you pick. Credentials are a separate matter: they go from an admin-only form into an encrypted vault, they are never shown in chat, and the agent only ever sees the field names — never the values.

What does it cost?

Priced per workspace, not per seat, with everyone in your company included. We are quoting during the private beta rather than publishing a price we would have to change in a quarter.

Can I sign up today?

Evalyst is in private beta, so access is granted rather than instant — you confirm your email, we come back to you. If someone at your company already has a workspace, we point you at theirs rather than starting you a second one, which is usually the faster route in.

Find out what your data says — and whether it's right.

Private beta. Tell us where to send the link and we'll take it from there.