User sessions on your Claude Connector ChatGPT App or MCP live inside their AI client, not your UI.
Armature shows you what users do with your product, and runs evals to make sure it keeps working.

Start for free
Backed by Y Combinator
1
MCP Analytics

See what users do.

list_customers
create_subscription
send_invoice402
send_invoice
search_customers
get_plans
list_invoices
update_plan
usage_report
Claude
User intent · 14:16 Set up a paid plan for ACME and invoice them
Agent thinking · 14:16 Invoice needs a billing contact — retrying with the owner email.
✓ Agent completed user task successfully 92score
With Armature
  • Session rebuilt
  • User intent captured
  • Agent thinking captured
  • Session evaluated
Use cases

Learn your users' top use cases

Armature's models read every session and identify what the user came to do. Sessions are grouped into use cases and ranked by volume and success rate, including the use cases you don't support yet.

Create & send invoices38%
Reconcile payments22%
Bulk refunds not supported yet14%
Export revenue report9%
Issues

Spot the top issues and fix them

Models scan every session for failures, loops and dead ends, group them by root cause and rank them by how many users encountered them, even when every API response was 200 OK.

127 Agent loops on missing auth scope↑ 43 this week fix →
64 Search misses "refund" phrasing
31 Export truncated by pagination
12 Rate limit hit on bulk updates
Session replay

Replay and investigate any session

Every session is scored on whether the user got what they asked for. Open any of them and replay the full trace: the ask, the agent's thinking and every call. When something breaks, you see exactly where.

User intentSet up a paid plan for ACME and invoice them
list_customers
Agent thinkingInvoice needs a billing contact — retrying with the owner email.
send_invoice402
send_invoice
✓ Agent completed user task successfully92score

Set up in a few minutes

Create your key

Sign up, generate an API key and drop it in your deployment secrets.

Prompt your coding agent

Paste one prompt and let your agent wire the SDK in.

Deploy

Ship it and watch sessions flow in within minutes.

Your users' data stays safe

Sessions can carry personal information and secrets, so we treat every one of them as sensitive. Detection models scan sessions and redact PII and secrets by default, before anything reaches storage.

2
MCP & CLI Evals

Turn the workflows your users actually run into evals

Two ways to create evals
Drafted by MCP Analytics

Use it, and your top use cases and issues arrive as ready evals.

Or written by you

Type the ask a user would send, and Armature drafts the rest. No analytics needed.

  • Real agents run each eval end to end, on all major models and harnesses.
  • A judge scores every run against your criteria.
  • Catch regressions and prove improvements before you ship.
Eval suiterunning
Analytics + evals

Better together.

MCP Analytics sees real usage
MCP & CLI Evals proves it keeps working

Each works on its own. Together they close the loop.

Start for free

Or reach out to us for custom needs.

Free $0
  • Session analytics and evals New1,000 credits / month: 1,000 sessions, 100 eval runs, or any mix. Then $50 per 1,000 extra credits.
  • Unlimited projects launch offer
  • Unlimited users
  • 7-day retention
  • Dashboard access
  • Session replay
  • Use-case grouping launch offer
  • Issue identification launch offer
Start for free No credit card required
Custom Let's talk
  • Everything in Free
  • Priority support & SLA
  • SSO / SAML
  • Audit logs
  • Custom retention
  • Dedicated onboarding
Contact us

Common questions

01How does Armature capture sessions?
You add our SDK to your MCP server, Claude Connector or ChatGPT App backend. It takes a few lines and one deploy. From there Armature rebuilds every session, classifies it and scores it automatically.
02Do I have to change my MCP server?
No. You keep the server you already run. The SDK wraps it to capture sessions and changes nothing about how it behaves for your users.
03Which agents and clients do you support?
Any client your users bring: Claude, ChatGPT, Cursor, Codex, Gemini CLI and the rest. If it can reach your MCP server, Armature can capture the session.
04Can Armature run evals on my MCP too?
Yes. Alongside analytics, Armature runs evals against your MCP server: real agents replay the asks your users send, across models and harnesses, and a judge scores every run against pass criteria you define. You can generate evals from your top use cases and issues, or write them from scratch.
05How is this different from PostHog, Amplitude or Mixpanel?
Those tools track humans clicking your UI. When a user delegates to an agent, the session happens inside Claude or ChatGPT and your UI never sees it. Armature captures exactly those sessions.
06How is this different from LangSmith or Langfuse?
They observe the agents you build yourself, and they serve the engineers who build them. Armature shows product teams how your users' agents experience what you ship.
07Is my users' data safe?
Yes. Detection models scan every session and redact PII and secrets by default, before anything reaches storage. You control retention and can delete your data at any time. Ask us for the full security details, we are happy to share them.
08What does it cost?
It is free up to 1,000 credits a month: 1,000 analytics sessions, 100 eval runs, or any mix, with unlimited projects and users. Past that you pay $50 per 1,000 credits. Custom plans cover teams with serious traffic.

Start improving your agent experience

Capture the sessions users have with your product through any AI agent. See what they ask, watch how agents deliver and fix the issues they encounter. Then run evals that catch regressions before you ship.

Start for free