KeyKit

You got the key. You don't have the week.

KeyKit does the testing you don't have time for and tells you whether an API actually works for your requirements. Fit, Partial Fit, or No Fit.

Built by Mulberry, from a decade on the buyer side of these contracts.

Start an assessmentTalk to us

Don't have the week either? Have Mulberry run the evaluation with you.

A trial key doesn't tell you if it works for you.

You can see in a demo that an API returns data. You can't see whether it returns the coverage, freshness, and fields your use case needs, at the volume you'll actually run. Finding that out properly takes time you don't have, a testing setup you'd have to build, and enough experience to know what to check and what good looks like. So most buyers skip it, sign on the demo, and find the gaps after they've paid.

Don't have a provider shortlisted yet? Start with the market map on Sourced

How it works

The evaluation you'd have to build yourself, already built.

It runs for you.

Bring a trial key. KeyKit does the evaluation, no engineering build, no week of setup. Simple mode gives a plain-English read, Advanced mode gives full framework control, switch any time.

It tests everything that matters.

Every check your requirements call for, run the same way on every provider, so nothing important gets skipped and two providers are compared on the same terms.

It knows what the results mean.

Which numbers are strong, which are a red flag for your use case, and what to trust, built by Mulberry from a decade of signing these deals.

Evaluation #007
COMPLETE
Fit
Historical Depth
38 mo · req. ≥ 24 mo
92
Fit
Freshness Lag
1.8 hr avg · req. ≤ 4 hr
88
Partial
Field Completeness
82% full · req. ≥ 90%
61
Fit
Deduplication Rate
1.1% dupe · req. ≤ 5%
94
No Fit
Rate Limit
800 req/hr · req. ≥ 5,000
18
Fit
Response Latency
340 ms p95 · req. ≤ 800 ms
85
4 Fit
1 Partial Fit
1 No Fit
Evaluation frameworks

The questions a trial key can't answer on its own

Will it hold up in production?

Latency, rate limits, and availability under real load, not demo conditions.

Is it as fresh as they claim?

We measure actual ingestion lag against your tolerance.

Will it handle your real queries?

The boolean, nested, and messy queries your use case needs, not just the simple ones in the demo.

Is what you're seeing normal?

How your results compare across everyone testing the same provider.

30 structured tests sit behind these questions, across coverage, data quality, freshness, reliability, compliance, and more.

See all 30

Coverage

2 runs

Does the dataset cover the time range and regions your use case requires?

Historical DepthGeographic Coverage

Data Quality

4 runs

Are records complete, canonical, and free of duplicates before they hit your pipeline?

Field CompletenessDeduplication RateCross-Query ConsistencyProvenance Metadata

Determinism

4 runs

Does re-querying the same parameters return the same results? Critical for incremental pipelines.

Result Set StabilitySort Order StabilityCount StabilityField Value Stability

Freshness

2 runs

How stale is "live" data? We measure actual ingestion lag against your stated tolerance.

Freshness LagLag Distribution

Query Complexity

6 runs

Can the API handle the queries your use case actually needs, or only the simple ones in the demo?

Basic Keyword QueryBoolean LogicNested BooleanWildcard & FuzzyField-Scoped QueryComplex Multi-Clause

Scale & Reliability

3 runs

Performance and stability under realistic load, not cherry-picked conditions.

Response LatencyRate Limit DiscoveryAvailability Check

Language & Scripts

1 run

Does multilingual content arrive correctly encoded and attributed?

Language Coverage

Stress & Edge Cases

4 runs

What breaks at the edges? Edge-case testing surfaces failures before production does.

Malformed Query HandlingEmpty Result HandlingRate Limit BreachDeep Pagination

Compliance & Cost

3 runs

Is sensitive data scoped correctly? Does cost hold at volume?

PII / Sensitive Data ScanQuota AccountingAuth & Scope Boundaries

Benchmarking

1 run

Side-by-side scoring against your current provider or an alternative. Apples to apples.

Category Benchmark

Every result is Fit, Partial Fit, or No Fit against your requirements, not an industry average.

After you sign

The provider that looked great in the demo gets thinner every month.

Coverage drops. Latency creeps. Fields go missing. Health Checks re-run your evaluation on a schedule and flag it when quality drifts, so it is your finding, not your incident.

Weighing two or three options? Run the same evaluation across all of them and see them side by side.

Health CheckWeekly · Mondays
Jun 02
84
Jun 09
81
Jun 16
71
Jun 23
68
Jun 30
77
Score dropped 13 pts Jun 16 · Freshness lag exceeded threshold
TEST
PROVIDER A
PROVIDER B
Freshness Lagdiff
Fit88
Partial61
Field Completeness
Fit94
Fit90
Response Latencydiff
Fit85
No Fit22
Availability
Fit99
Fit97
Deduplication
Fit91
Fit88
Pricing

One plan, one price.

KeyKit is $250 a month, the same account for buyers and providers. Buyers evaluate any data or model API against their requirements. Providers get that same account, plus private self-testing and the ability to sponsor customers (see below).

KeyKit
$250/month

$2,500/year

Evaluate any data or model API against your requirements.

  • Unlimited evaluations
  • Simple and Advanced mode
    Switch between plain-English summaries and full framework-level control.
  • Health checks
    Run any evaluation on a recurring schedule. Daily, weekly, or monthly.
  • Evaluation comparison
    Compare any two completed evaluations side by side, framework by framework.
  • Full framework library
    30 evaluation frameworks across 10 groups.
  • Request logs and raw output
Start an assessment
For providers

Prospects don't say no. They go quiet.

A trial key hands the work to your prospect: find the time, wire it up, decide what the numbers mean. Most never get to it. Sponsor a KeyKit account instead and they get a neutral evaluation of your API against their own requirements, without the setup. If you're confident in your data, an independent verdict is the strongest thing you can put in front of them.

Sponsor a customer for 30, 60, or 90 days, at $125, $250, or $375 (half the monthly rate). They run their own evaluations and keep the account at $250 a month if they want it. No automatic charge.
It's KeyKit's neutral verdict, not a demo you control. They see the results, you don't. That is why they trust it.
Your own account also lets you test your endpoints privately, before a buyer ever does.
Shortens the sales cycle and sets you apart from providers still sending benchmark PDFs.
Book a meeting