Inclusion criteria
This page is written for the site it describes. The index went live on 23 August 2026; this page publishes alongside it, and its present-tense statements hold from the day it publishes.
This page is the whole of how EvalIndex decides what it lists. There is no other test, no editorial shortlist and no discretionary override.
A product is in one of five states here, and this page says which:
- Listed — it passed every rule in sections 1 and 2, and its record shows the evidence for each. There is one exception, and it is always marked on the record's own face: a listed record that stops meeting a rule keeps its place in the comparison table for up to 30 days while that mark stands (section 4). So a listed row can be a marked row, and the mark says which criterion it fails and since when.
- Queued — it is waiting its turn. Some queued products have already been assessed and passed every rule; others have not been assessed yet. Either way, being in the queue is not a finding against a product: a queued product has failed nothing.
- Refused — it failed a named rule before it was ever listed. The record of that assessment says which rule it failed and what we read to decide it.
- Delisted — it was listed, and a re-verification found that it no longer meets a criterion while the product itself is still alive and still sold. After the 30-day marked grace above it leaves the comparison table for the delisted list. Its record page stays, marked with the criterion and the date, and it returns to the table when it recovers (section 4).
- Discontinued — the product itself is gone: the vendor has said so, or the product's site or domain has failed the repeated checks in section 4. Delisted and discontinued are not the same finding, and we never publish one as the other.
Nothing on this page can be bought. Section 6 says that plainly, and section 7 says what happens when this page changes.
We apply these rules in the order they are written and use the same test for every product. For each rule in sections 1 and 2 the record names the page we read to decide that rule — its evidence link — and the date we read it. Where more than one page mattered to a rule, the record's notes name the others. Those source URLs are published on the record, so any reader can repeat the check we made.
1. What counts as an LLM observability platform
The index covers products whose own purpose is observing, tracing, evaluating or monitoring applications that call large language models. A product qualifies when all four of the following are true.
1.1 It is an obtainable product. A software product or hosted service — paid, free of charge, or open source — offered to other organisations or developers for use in their own applications. A research paper, a specification, a blog post's example code, a consultancy, an internal tool with no public release, and a benchmark or model-comparison leaderboard are not products for this purpose. A leaderboard evaluates models; this index covers products that observe the software you build with them.
1.2 It has LLM-specific features, documented as such. The product's own documentation contains a page, section or named feature dedicated to LLM or AI-application work, describing behaviour that is specific to model calls. At least one of these four capabilities must be documented:
- Trace capture — recording the prompts, model responses and the sequence of calls an application makes, as retrievable traces or spans.
- Evaluation — running scorers, graders, model-as-judge checks, human review or test datasets against recorded or newly generated model outputs.
- Monitoring on model-specific measures — token counts, cost per call, latency by model or provider, error or refusal rates, or output quality scores, reported per model call.
- Prompt or dataset management joined to any of 1–3 — versioning prompts or datasets and comparing recorded runs across versions.
A page that says LLM calls will appear as ordinary spans, with no LLM-specific field, view, scorer or setting documented, does not satisfy this rule.
1.3 The LLM capabilities have their own place in the documentation. At least one capability from 1.2 is documented in a section of the product's own documentation that is reachable from that documentation's top-level navigation, at the same level as the product's other documented features. A single integration page, a marketing landing page and a launch blog post are none of them a top-level documentation section.
What the product's home page says is not part of this test, in either direction. Marketing copy neither qualifies a product nor disqualifies one: this rule is decided by opening the documentation and looking at its navigation.
This is the boundary against general-purpose monitoring. An application performance monitoring, logging or error-tracking product whose only LLM material is a marketing landing page, a launch blog post or a single integration page — with no LLM-specific capability from 1.2 documented at the level 1.3 describes — does not qualify. A general-purpose product that does document such capabilities at that level does qualify: the test is what the documentation contains, not what the company is otherwise known for.
1.4 It observes your application, not only the vendor's own service. The documentation shows the product working with model providers, frameworks or SDKs other than the vendor's own hosted models, or shows it being self-hosted against them. A usage-and-billing dashboard that can only report on one vendor's own API is that vendor's dashboard, not an observability platform for the applications you write.
One record covers one product. Every record names one listed product — the hosted or primary product the record is about — and every finding on the record is a finding about that product and no other.
- Where a vendor offers products that have both a separate name and a separate code base, they are separate products. Each is assessed on its own evidence, each can be listed, queued or refused on its own, and they are never merged into one record.
- Where an open-source project and a hosted service share a name and a code base, they are one record, and the record states which findings apply to which.
- It follows that a dated release in a repository that is not the listed product's own code base is not evidence of the listed product's maintenance under 2.3, however closely the two are related. The evidence has to be about the product on the record.
2. Eligibility rules
Each rule below is decidable by a stranger from public evidence, without an account beyond a free self-serve signup, and without payment. Every rule is applied again in full at each re-verification (section 4), not only when a product is first assessed.
2.1 Self-serve obtainable. From the product's own website a reader can obtain the product without speaking to anyone at the vendor: an account they can create themselves, a download, an entry in a public package registry, or a public source repository. A product whose only route is a sales conversation — a "contact us for access", "request access" or "book a demo" form and nothing else — fails this rule, and the record of the assessment says so. This is a deliberate line: the index compares products a reader can go and get today.
What a registry entry or a repository has to be. A public package registry entry or a public source repository satisfies this rule only where what it distributes is the listed product — something the reader can run for themselves. A client library or an instrumentation SDK that sends data to a hosted service is not that hosted service: publishing one does not by itself make the hosted product obtainable. Where the hosted product's only other route is a sales conversation, it fails this rule however many public packages the vendor ships, and the record names the packages we looked at and says why they did not settle it.
2.2 Public documentation. Product documentation is readable without an account and without payment, at a URL linked from the product's own site, and it documents at least one capability from 1.2. Documentation available only after a sales call fails this rule.
2.3 Active maintenance, dated. At least one of the following carries a date no more than 12 months older than the assessment date, and that date is visible to any reader:
- a release or version tag published by the vendor — release notes, a changelog entry, or a dated release in a public package registry;
- a dated product-update or changelog post on the vendor's own site;
- a dated commit in the product's public source repository, where it has one;
- a documentation page carrying a visible last-updated date.
Date precision. A date given only to the month counts, and we read it as the last day of that month for this window — that is the reading that cannot make a maintained product look unmaintained, and it is applied the same way to every product. A date rendered without a year is not evidence on its own; it counts only when a tag, changelog entry or other artefact carrying the year corroborates it, and then the record cites that corroborating artefact as the evidence. We do not assume any platform's convention about when it does or does not render a year.
An undated page is not evidence for this rule. If nothing dated inside the window can be found, the product is not eligible, and the record of the assessment says which sources we looked at.
This rule is tested again on the publication date (section 3).
2.4 A working product website. The product's own site loads over HTTPS and serves the product's own content. A parked domain, a holding page, a certificate error that blocks access, or a domain that now redirects to an unrelated site fails this rule.
One failed load never decides this rule, anywhere on this page. At assessment, a product fails 2.4 only when the site is in one of those states on two checks made at least 3 days apart. For a record that is already listed, the longer flow in section 4 governs — three checks made at least 7 days apart, so at least 14 days. The bar is higher for a listed record on purpose: delaying an assessment costs a vendor a few days, and removing a published record over a transient outage would be wrong in a way we could not undo.
This rule is tested again on the publication date (section 3).
2.5 Enough public evidence to assess. Every rule in sections 1 and 2 must be decidable from public sources, and the record's identifying fields must be establishable from them: product name, vendor, product URL, documentation URL, deployment options (hosted, self-hosted, or both), and at least one capability from 1.2. If any of those cannot be established publicly, the product is not listed until it can be, and we say which one was missing.
The comparison fields are a different matter, and this is what a reader will see. Where a vendor genuinely does not publish a figure — data retention, data residency, a price — the record shows that field as unstated, names the vendor page we checked for it, and gives the date we checked. Where it helps the reader, the record also says what the vendor publishes instead — that it prices per request rather than per trace, say — because a reader comparing prices is better served by that fact than by a blank. What the record never does is fill the cell with an estimate of the figure the vendor did not publish. "Unstated" is a fact about what the vendor publishes, not a mark against the product, and it never affects eligibility. Publishing a price is not a criterion; nor is publishing a retention period.
What is not a rule. Nothing here turns on a product's popularity, size, funding, age, download counts, social following, whether its vendor replies to us, or whether its vendor has bought anything. Those facts are not part of the test and are not used to decide what is listed.
3. The launch set, the queue, and how the index grows
The launch cut, as a historical fact. EvalIndex assessed 15 products on 22 August 2026, and the site went live with the launch index on 23 August 2026. Thirteen of the fifteen met every rule in sections 1 and 2, and twelve of those thirteen were the records published on that second date. Two of the fifteen did not, and the state model at the top of this page requires us to say which and why:
- Fiddler AI — refused under criterion 2.1. On the vendor's pricing page every tier, the free one included, routes to a "Request demo" form; we found no self-serve account creation on the vendor's site.
- Traceloop — refused under criterion 2.3. We found no dated release, changelog entry or last-updated date for the hosted platform itself; the dated changelog we found belongs to OpenLLMetry, a different code base, which the last bullet of section 1 excludes as evidence of the listed product's maintenance.
Each of those records names what we read and the date we read it, and either product is assessed again on request or on new public evidence. A refusal is a finding about the public evidence on that date, not a verdict on the product's quality.
Of the thirteen eligible, twelve were published at launch and one — W&B Weave — entered the queue. The twelve were chosen by alphabetical order of product name, for one reason stated plainly: twelve complete records was as much as this index could research and stand behind at once, and alphabetical order is the one cut no product can influence and every reader can check. The thirteenth is queued, not refused: it passed every rule, and a place in the queue is a position, never a finding. The assessed set is published with the date it was assessed, so the cut can be checked against the rule.
That was a one-time cut. Alphabetical order has governed nothing since, and it does not govern the queue.
The queue is first in, first out. A product's place in the queue is set by the date it entered the queue: the date we received the request, or, for a product we found ourselves, the date of the assessment sweep that found it. Earlier date, earlier place. Where two products entered on the same date, alphabetical order of product name breaks the tie — and where their names also reduce to the same sort key, the final tie-break under How names are sorted below decides. Nothing else moves a product up or down, with the single exception in Expedited assessment below. The queue is published with each waiting product's entry date, so anyone can check the order against this rule.
How fast the queue moves. We aim to publish one new record a week and to re-verify 12 records a week (section 4). Those are aims, not promises: some weeks will be more and some fewer. A record that takes longer to research publishes late rather than thin — there are no placeholder rows and no stub pages. The queue page shows the date each record was actually published and the date each waiting product entered, so the real rate is visible rather than asserted.
How large the index gets. The index is capped at 40 records, and the cap is set with margin on purpose. At 12 re-verifications a week, 40 records are cycled in 40 ÷ 12 = 3.33 weeks — about 3 weeks and 2 days — comfortably inside the 45-day staleness threshold in section 4. The largest number that cadence could in theory keep inside 45 days is about 77 records (45 ÷ 7 × 12 = 77.1). Forty is well below that, so that a missed sweep or a slow week does not push records past the threshold. If the cadence or the cap changes, it changes here first, with the reason, under section 7.
When the index is at 40 records, eligible products stay in the queue until a slot opens; a record leaving the comparison table under section 4 opens one.
Tested again before publication. A product's assessment can sit in the queue for a long time, so 2.3 and 2.4 are tested again on the publication date, and a record whose last full verification is more than 45 days old on its publication date is verified again in full before it publishes. Nothing publishes already stale.
How names are sorted. This is the listing order of the comparison table, and the tie-break in the queue.
The name we sort on is the product's name as the vendor's own site writes it on the date we assessed the product. We resolve it by this ladder, in order, and the record says which step decided it and shows what we read:
- The home-page masthead, where the masthead names the product. Where the masthead and the page title disagree and both name the product, the masthead is the name we take, and the record shows both.
- The product name in the home-page title, where the masthead is a company mark or a tagline rather than a product name. A tagline cannot name a product, so the title's product name stands.
- The vendor's own product surfaces — the product navigation and the pricing page — where the home page names the company and those surfaces name the listed product distinctly. The name those surfaces use is the one we take.
A transition or acquisition banner does not rename a product. A home-page announcement that a product is becoming, joining or being renamed to something else is a statement about the future; the name changes here when the vendor's own masthead or title changes, and not before.
Every other name the product carries — a former name, an acquiring company's name for it, a code-base name, and any other name the vendor also uses — is published on the record as an alias, so a reader looking under any of them finds the record.
A product that is renamed keeps its record and its URL. The record notes the new name, the date we saw it and the name it replaces, and the record re-sorts under the new name at its next re-verification, not before — so the table never reorders itself between one visit and the next without a dated note on the record saying why.
The comparison itself: case-insensitively, character by character, with
digits ordered before letters. Every character that is not a letter or a
digit — spaces, punctuation, &, +, #, and anything else — is
removed before comparing. An initial "The", "A" or "An" is kept and
sorted as written. If two names reduce to the same key, the earlier
queue-entry date sorts first. If the entry date is the same as well, the
earlier request timestamp sorts first. Where two requests are genuinely
simultaneous, neither waits on the other: both records publish in the
same week, and the queue page says that is why.
Expedited assessment. Queue position can be expedited for a fee; the terms are on the paid placement policy page. What the fee does, exactly: it moves a product to the front of the queue of products not yet published. What it does not do: it does not affect eligibility, whether a product is listed at all, what the record says, where the product appears in the table or in any listing order, or the verdict — and an expedited assessment is published in full whatever it finds, including when what it finds is unflattering. A product that fails any rule in section 2 is not listed however its assessment was ordered; the paid placement policy page says what happens in that case.
Expedited assessment is offered only while there is more room than queue. It is offered while the number of free slots — 40 less the records in the comparison table — is greater than the number of products waiting in the queue. Whenever free slots are fewer than or equal to the queue length, no expedited fee is offered and none is accepted; that includes the index being full at 40, where there are no free slots at all. Any expedited fee already paid for an assessment that cannot be published on those terms is refunded in full.
The reason is one line: at that point queue position would decide whether a product is published, not only when — and a fee that decided that would be a fee for presence in the index. We would rather lose the sale than sell presence.
4. Freshness, re-verification, marking and removal
Every record carries the date it was last verified in full, and its age alongside that date.
Re-verification. We aim to re-verify 6 records per sweep, twice a week — 12 records a week. A re-verification re-reads the sources named on the record, updates the fields that changed with the change shown on the record, and re-applies every rule in sections 1 and 2 in full. A record's last-verified date moves only when its sources have actually been re-read. We never re-stamp a record to make it look current.
A listed record that stops meeting a rule. Eligibility is not an entrance test with no exit. Where a re-verification finds that a listed record no longer meets a criterion, the record says so on its face — "no longer meets criterion 2.3 as of 14 September 2026" — with what we checked and when. The record stays fully readable and stays marked. If it still fails 30 days after it was marked, it is delisted: it leaves the comparison table for the delisted list, its record page stays at the same URL, and that page states the criterion it fails, the date it was marked and the date it was delisted. If it recovers first — a new dated release, the site back up, documentation public again — the mark comes off and the record says when it came off and why. A delisted record returns to the comparison table at the next re-verification that finds every rule met again, and the record says so with its date.
Delisting is not discontinuation. A delisted product is alive and still sold; it stopped meeting a criterion. The discontinued list below is reserved for products that are genuinely gone, on the checks set out there. We do not publish a trading product as discontinued.
The asymmetry, stated on purpose. A newcomer and a listed record in the same public condition do not get the same answer, and the difference is deliberate. A newcomer that fails a rule is refused before it is ever listed. A listed record that fails the same rule keeps its place for at most 30 days, marked on its face, because readers already hold links to it and a row that vanishes without notice tells them nothing. The grace is in the timing only: the mark is the same public fact either way, it names the same criterion, and at the end of the 30 days the record leaves the table.
Stale records. A record whose last-verified date is more than 45 days old is labelled stale, on its own page and in the index, showing the date and how old it is. The data stays fully readable; the label states its age. A record too stale to stand behind is marked stale, not quietly left looking fresh.
Removal and discontinuation. A record is marked discontinued when any of the following is true:
- the vendor publishes that the product is discontinued, end-of-life, or no longer sold;
- the product's website or its documentation fails to load on three checks made at least 7 days apart;
- the product's domain no longer serves the product — parked, offered for sale, or redirecting to an unrelated site — on three checks made at least 7 days apart.
The last two are the checks 2.4 refers to for a record that is already listed.
A discontinued record keeps its URL and states the date it was marked, what we checked, and when. It leaves the comparison table and moves to the discontinued list. We do not delete records silently: an index that erases its own history cannot be checked by anyone. A record is deleted outright only where we are legally required to remove it, and the URL then says that the record was removed and on what date.
Corrections. Anyone can tell us a record is wrong, at
hello@evalindex.dev. We re-check that field against public sources,
fix what is wrong, and note the correction on the record with its date.
A correction updates the checked date of the field we re-read, and only that field. It does not move the record's last-verified date: that date means the whole record was re-read, and it moves only on a full re-verification. One corrected field cannot make eight un-re-read fields look freshly checked.
A correction changes what the record says; it does not change whether the product is listed, which is decided only by sections 1 and 2.
5. Requesting a listing
Anyone — a vendor, a user of one, or a reader who thinks something is
missing — can ask for a product to be assessed. Write to
hello@evalindex.dev with the product name and its URL. Links to the
documentation and to dated release or changelog pages help, because they
are the evidence rules 2.2 and 2.3 are decided on.
The rules in sections 1 and 2 apply to a requested product exactly as they apply to one we found ourselves. Requesting changes nothing about the assessment: not its result, not what the record says, not where the product appears. A requested product joins the queue at the date we received the request, under section 3; it does not go to the front.
We answer requests sent to hello@evalindex.dev, normally within a few
days. The answer says either that the product meets the criteria and
where it enters the queue, or which rule it fails and what public
evidence would settle it. No fee is required to be assessed or
listed.
6. Payment, and what it cannot do
- No fee is required to be assessed or listed. No product is in this index because it paid, and none is left out because it did not.
- No payment can prevent, alter or remove an assessment, change a verdict, change what a record says, or change where a product appears in the table or in any listing order. An expedited assessment moves each unexpedited request that was ahead of it one place back in the queue; requests behind it are unaffected. It never changes what any assessment says, and it is not offered at all once free slots are fewer than or equal to the queue length (section 3).
- The two things that can be bought are on the paid placement policy page: an expedited assessment, which moves queue position and nothing else (section 3), and a vendor-authored block, which adds the vendor's own words to their record, below our independent assessment, which it does not alter. Every vendor-authored block carries this label inside the block, above the vendor's words: Vendor-authored · paid feature · published <date> · unedited. The label says that it was paid for, because that is the part a reader needs to know.
- There are no paid placements, and none are sold today. The two things above are the whole of what can be bought. If measured traffic ever justified sponsor slots, they would be sold as a labelled band outside the index table — never rows in it, holding no position in any listing order — and this page would list them here as a third purchasable thing before any were sold.
- The index has no user reviews and no star ratings, paid or unpaid. Every assessment on this site is our own first-hand work, with its sources linked.
- We do not take payment to soften or withdraw a finding. An unflattering assessment is published as found.
7. How these criteria change
These rules can change. Every change is published here with its date, and the wording it replaced is kept below, verbatim.
A change is never applied to one product and not another. When a rule changes, every existing record is tested against the new rule at its next re-verification, and a record that then fails is marked and handled under section 4. We do not quietly re-label past assessments to fit a new rule, and we do not leave a new rule applying only to newcomers.
One rule about the rules: where a criterion turns out to be ambiguous, we fix it by rewriting it here, for everyone, before applying it — never by deciding one case one way and leaving the wording as it was.
Changes to this page
| Version | Date | What changed |
|---|---|---|
| 1 | 22 August 2026 | First drafted. |
| 2 | 22 August 2026 | Rewritten before first publication: launch-cut honesty, ceiling/fee interaction, rename rule, FIFO queue, obtainability tightened, eligibility re-testing. |
| 3 | 22 August 2026 | Cured before first publication: launch record corrected to the assessment as actually made (13 eligible, 2 refused on named rules, 12 published, 1 queued); the 2.1 registry/repository test, the name-resolution ladder and the delisted state published as rule text; the expedited-fee carve-out re-triggered on free slots against queue length; the fairness asymmetry, the paid-placement inventory, the evidence-link promise and the final tie-break stated as they actually work. |
| 4 | 23 August 2026 | Cured before first publication: the launch dates stated exactly (assessment 22 August 2026; site live 23 August 2026). |
Superseded wording
Each time a rule on this page changes, the wording it replaced is kept here in full, under the version that replaced it, so that anyone can read exactly what the rule used to say alongside what it says now.
Versions 2, 3 and 4. Each of those three versions replaced an earlier draft — version 1, then version 2, then version 3 — and every one of those drafts was written and replaced before this page was ever published. None of them was ever in force, so none of them has wording to supersede: no rule text other than the text above this table has ever governed a single assessment. Version 4 is the wording this page publishes with; from that day on, every rule this page changes has its previous wording kept here in full, under the version that replaced it.