The census
Laboratory evidence that is public, and effectively invisible.
This is the one result on this site with nothing to do with my variant. It is a measurement of a database, and it applies to every gene in it.
The short version
When a laboratory measures what a genetic variant actually does to a protein, it can deposit that measurement in ClinVar, the public database clinicians and genetic counsellors rely on. That evidence is real, it is expensive to produce, and it is filed correctly.
It is also absent from the compact version of the record that automated tools request first, and none of the tools surveyed reads it from the full version either. Several already have the full record on disk and simply never look at the field.
Of 12 widely used variant-interpretation resources, 11 have a determinable route to ClinVar and 0 of them read the functional-evidence field. The twelfth has no dedicated client, so the question does not apply to it.
That list includes tools that ingest the complete release rather than the compact summary. The data is on their disk. Nothing parses it.
The size of the difference
For one variant, the compact record runs 3,763 characters and the full record 21,611. The functional effect, the numeric result, the severity call and the evidence code exist only in the larger one. The compact format declares 74 elements and not one of them is functional.
The scan itself was a single pass over the 2026-06-27 release, 4,531,457 records and 5.82 GB, in 53 minutes on one machine.
How many variants this affects
variants across 622 genes carry functional evidence deposited by a laboratory other than the one bulk depositor.
of those, 5,359 variants, still have no confident clinical classification.
And this is not raw data waiting to be interpreted. 1,892 of them, 24.7 percent, carry a formal evidence code in the deposit text, meaning the laboratory that did the measuring has already stated how strong it considers its own result to be. The most common are the codes for functional evidence supporting a harmful reading, 1,110 times, and a harmless one, 1,106 times.
An earlier build of the data file reported a larger pool. The exclusion filter for the bulk depositor was tested against a shortened, four-name copy of each record’s submitter list rather than the real one, so records where that depositor sorted fifth or later slipped through. That accounted for 911 records exactly, and a second defect readmitted 15 more. Neither figure was a population of anything. The paper’s published numbers never depended on the faulty file, and the full working is kept as a companion note rather than the file being swapped out quietly.
The correction that has to travel with any count
One laboratory accounts for the overwhelming majority of records carrying this evidence, through a single very large bulk submission. Quoting the raw total without that correction would treat one submission as though it were hundreds of thousands of independent laboratory characterisations, and would misrepresent the field badly.
Every version of this analysis excludes that depositor before reporting anything. It is stated first because it is the most likely way the result could mislead.
Why unclassified does not mean neglected
This is the part that makes the finding more interesting rather than less. A classification in ClinVar requires that a variant has been seen in a person and submitted with clinical context. Functional evidence alone does not trigger one, and it is not supposed to.
So many of these are unclassified because no carrier has been reported yet. The evidence has been banked ahead of the first patient, sitting in a field the ordinary programmatic route does not return. Whether it is used when someone does turn up depends on whether the person interpreting the variant knows to look in the full record.
Two variants in the same small region of the gene I was studying carry the strongest functional evidence tier the framework allows, and have never been classified by any submitter. Both are absent from population data. That is evidence banked ahead of the first carrier, and noticing it is what prompted this census. The census then showed it is not a peculiarity of that region or that gene.
What this is not
It is a measurement of a database, not an audit of clinical practice. It shows the field is absent from one access route and unread by the tools surveyed. It does not measure how many real interpretation decisions were affected, which would require knowing which route each tool took in each case.
The omission is documented rather than hidden. The database publishes the behaviour in its own documentation. The finding is that the omission is consequential, not that it is concealed, and the result survives that correction.