brugada.net
Preprint, not peer reviewed. Posted publicly before review so that the reasoning and any errors are both visible. Treat every claim as provisional. Plain markdown source.

Version of record: 10.5281/zenodo.21799868, published 5 August 2026. That identifier is the citable address for this paper and it resolves at https://doi.org/10.5281/zenodo.21799868. It is a version identifier; Zenodo minted a second one that resolves to all versions, and the version identifier is the one to cite.

What this file is. This project's authoritative copy of the manuscript, SUBMIT_THESE/papers/PUBLISH_9_NTD_VUS_RESOURCE.md, which is the file the deposited PDF was built from. Synced 7 August 2026 by scripts/sync-manuscripts.mjs, which copies the source byte for byte and prepends this note. Nothing in the manuscript below has been rewritten for the website.

This text matches the version of record, on the evidence available here. The authoritative markdown was last written on 4 August 2026, before the submission PDFs were built that evening and before the deposit the following day, so nothing in it postdates the identifier above.

The limit of that claim, stated rather than hidden. It is an argument from timestamps, and for this paper it is still only that. (This paragraph used to say that no published PDF had ever been downloaded and no published text read back. That was true when it was written and stopped being true on 6 August 2026, so it is corrected here rather than left standing.) The published files of the divergent papers were downloaded from the Zenodo API on 6 August 2026 and diffed against the local text, which is what turned that divergence from an inference into a measurement. This paper was not among them. It is one of the copies believed unchanged on the strength of local modification times, which is the same class of inference the download exercise was run to replace. SUBMIT_THESE/ZENODO_DIVERGENCE_20260806.md records what was checked, and SUBMIT_THESE/V2_STAGING/evidence/DIVERGENCE_VERIFIED_20260806.md records what was not.

A separate limit, and it is not about this text. Six of the ten papers name deposited tables that the archive does not hold, in whole or in part. Paper 10's case was the worst and has been dealt with; the others have not. Do not assume this paper's data availability statement is true of the archive without checking it. The audit is SUBMIT_THESE/PAPER_10_DATA_STATEMENT_FIX.md.

None of this is peer reviewed, and none of it has been through a wet lab. No cell has been edited and no current has been recorded for this variant by this project. Every therapeutic statement in the manuscript below is a prediction.


A stability predictor with known blind spots nominates seventeen uncertain SCN5A N-terminal variants for testing

Ethan Bradley

Independent researcher, no institutional affiliation

ORCID: 0009-0008-8925-7975

Interpretation correction, 7 September 2026. The 0.79 kcal/mol cutoff is an exploratory nomination threshold derived from the reported antisymmetry discrepancy, not a calibrated Nav1.5 noise floor, predictive-error bound or statistical significance level. Opposite prediction errors can cancel in an antisymmetry sum, so that discrepancy cannot establish that smaller scores contain no information. The 17 nominations remain exploratory priorities; their positive predictive value is not established, and the 114 non-nominations do not establish benign status. Functional loss and folding destabilization are different endpoints: failure to nominate a functionally abnormal variant is not by itself a measured stability-prediction error. Agreement across Rosetta and OpenMM preparations assesses preparation sensitivity using the same predictor, not independent predictive validation. Neither nomination rule supplies per-variant null probabilities or an estimated false-discovery fraction. This correction supersedes the historical “measured noise floor,” “own measured error,” “no information,” and nominal-threshold chance-discovery interpretations below; reported scores, thresholds, lists and counts are preserved.

Abstract

Of 131 ClinVar variants of uncertain significance in the Nav1.5 (SCN5A) N-terminal domain placed on structure 8VYJ, 17 exceed an exploratory ThermoMPNN nomination threshold of 0.79 kcal/mol in both Rosetta- and OpenMM-relaxed preparations. The threshold derives from a reported antisymmetry discrepancy, not a calibrated Nav1.5 predictive-error bound or significance level. Three of four known loss-of-function benchmarks, R104W, R121W and Y87C, were predicted neutral or stabilizing (PMID 22739120, PMID 32815768), while the ER-retained variant A124D received strongly destabilizing predictions (PMID 22529811). These observations motivate caution but do not establish positive predictive value; functional loss and folding destabilization are different endpoints. The 114 non-nominations do not establish benign status. The 17 nominations are grouped into three exploratory tiers by substitution chemistry and structural context. A separate nine-paralogue conservation analysis identifies 41 of 141 uncertain variants at positions as variable as two benign controls; this comparison is descriptive, not independently calibrated clinical evidence. Overlap between the calibration gate and named validation examples prevents an independent-validation claim. Neither list has an estimated false-discovery fraction. Both are exploratory resources for selecting functional experiments, not clinical classifiers.

A key to the terms used here

Why this list exists and what it cannot do

The Nav1.5 N-terminal domain carries a large number of ClinVar variants of uncertain significance and almost no published functional data to resolve them. A laboratory with a patch-clamp rig has to choose where to start. I built this list to give that choice some structure, and I built it after first measuring, on the same domain, what folding-stability prediction can and cannot see.

The companion analysis reports neutral or stabilizing predictions for three variants with published dominant-negative effects, R104W, R121W and Y87C (PMID 22739120, PMID 32815768), strongly destabilizing predictions for the ER-retained variant A124D (PMID 22529811), and near-zero scores for two benign controls. These are comparisons with functional or clinical labels, not independently measured folding-stability endpoints. They do not establish predictive sensitivity, specificity or positive predictive value for this domain.

The reported benchmark results do not establish a clinically interpretable ranking of all 131 uncertain variants. I therefore report an exploratory nomination list using the same positive threshold in both structural preparations. This selection rule does not establish positive predictive value or imply that scores below the threshold contain no information.

How the list was built

Structure and predictor

Variants were scored against PDB entry 8VYJ, chain A, a cryo-EM structure at approximately 3.6 Å resolution with associated map EMD-43662. Stability changes were predicted with ThermoMPNN (Kuhlman Lab), using the tool's bundled default checkpoint. No separate version number for the checkpoint is recorded in my working notes beyond its default label. Structures were relaxed by two independent protocols, Rosetta and OpenMM, before scoring; specific software version numbers for these relaxation runs are not recorded in my working notes.

Variant set and filtering

I started from 172 ClinVar missense records in residues 12 to 130. These were retrieved through NCBI E-utilities; the exact query string and the retrieval date are not recorded in my working notes, and I state that plainly rather than guess at either. From the 172, I removed 9 nonsense or stop-gain records, since these are not a missense stability question, and 12 records at positions unmodelled in 8VYJ, an unresolved gap spanning residues 38 to 48. That leaves 151 scorable variants. Every one of the 151 wild-type residues was checked against the 8VYJ sequence directly, and all 151 matched, with zero mismatches. Of the 151, 131 are variants of uncertain significance and 20 already carry a ClinVar classification; the historical account describes the 20 as withheld, but the overlap with the calibration gate prevents treating the later comparison as independent validation.

Scoring and the calibration gate

The historical account states that each of 151 scorable variants was scored across 24 relaxed structures, 12 from the Rosetta preparation and 12 from the OpenMM preparation, and reports 3,576 individual predictions. These counts do not reconcile: 151 × 24 is 3,624, a difference of 48 predictions. The original run manifest is needed to determine whether evaluations were missing or a reported count was incorrect; neither explanation is selected here. The nomination rule requires the mean score in each preparation, taken separately, to exceed 0.79 kcal/mol.

Before any uncertain variant was scored, I ran a gate on seven variants with known status:

Variant Class Rosetta prep OpenMM prep Gate role
A124D ER-retained (PMID 22529811) +2.154 ± 0.014 +2.389 ± 0.068 exceeds threshold: PASS
R104Q proband's variant, functionally characterized +1.427 ± 0.006 +1.451 ± 0.254 reference point only
R104W dominant-negative (PMID 22739120) -0.015 ± 0.012 +0.627 ± 0.178 benchmark non-nomination
R121W dominant-negative (PMID 22739120), contested classification (PMID 32815768) -1.263 ± 0.021 -0.427 ± 0.519 benchmark non-nomination
Y87C dominant-negative (PMID 32815768) -0.030 ± 0.003 -0.119 ± 0.055 benchmark non-nomination
R34C benign control (PMID 11960580) -0.016 ± 0.024 +0.028 ± 0.026 must stay near zero: PASS
V125L benign control -0.258 ± 0.014 -0.197 ± 0.156 must stay near zero: PASS

A124D exceeds the threshold in both preparations, both benign controls stay within it, and A124D separates from the worst-scoring benign control by 2.17 kcal/mol in the Rosetta preparation and 2.36 kcal/mol in the OpenMM preparation. All seven values reproduce prior recorded results from this pipeline to within 0.001 kcal/mol, the strongest check available that the scoring run was rebuilt correctly. I also verified the sign convention independently, scoring the well-characterized cavity-creating variant p53 Y220C on structure 2OCJ: the raw output was +1.823 kcal/mol, confirming that the model reports destabilizing substitutions as positive and that no sign flip was applied anywhere in this analysis.

Exploratory threshold and its reported basis

The nomination threshold is 0.79 kcal/mol. In the reported analysis of 144 mutation pairs in soluble proteins unrelated to Nav1.5, predicted ΔΔG(A to B) plus ΔΔG(B to A), which would be zero under exact antisymmetry, averaged -0.79 ± 0.77 kcal/mol. This measures internal inconsistency, not predictive error: opposite errors can cancel. It therefore does not establish a Nav1.5 noise floor or show that smaller scores are uninformative. Separately, the reported Ssym benchmark has a predicted standard deviation of 0.90 against a true standard deviation of 1.54. Scores here are exploratory predictor outputs, not calibrated absolute folding free energies.

The historical account states that the threshold was fixed before uncertain variants were scored. The reported nomination counts at alternative thresholds are retained below; they do not calibrate decision errors.

Threshold (kcal/mol) VUS nominated Benign controls nominated
0.50 30 0
0.79 (used) 17 0
1.00 11 0
1.25 7 0
1.50 5 0
2.00 1 0

Neither benign control is nominated at the reported thresholds. This small, overlapping control set does not establish generalization specificity.

Relative solvent accessibility, reported below for context, was computed with mkdssp (version not recorded in my working notes) using the Tien et al. (2013) empirical maximum-accessibility scale. This reproduces prior values for R104 closely, 11.7% against a prior 11.6%, and for D82, 8.6% against a prior 8.5%, but differs for D84, 18.2% against a prior 16.0%. I report that discrepancy rather than harmonize it away. RSA is descriptive context only and was never used as a nomination criterion, so this discrepancy does not change any result below.

Seventeen variants exceed the threshold in both structure preparations

Of the 131 scorable uncertain variants, 17 exceed 0.79 kcal/mol in both preparations and 114 do not. Non-nomination does not establish tolerance. The benchmark misses motivate caution, but they do not show that every non-nomination contains zero information.

The 17 fall into three exploratory tiers based on substitution chemistry and structural context. Similarity to the buried-charge substitution A124D motivates follow-up; it does not establish predictive accuracy.

Tier 1, matching the chemistry of the A124D benchmark.

Variant ClinVar Rosetta OpenMM RSA Charge change
N97K VCV002989748 +1.652 +1.504 13.4% +1
A107D VCV001371233 +1.358 +1.221 6.6% -1

Both introduce a charge at a buried site, the same chemistry as A124D. N97K is the strongest single nomination in the set, and it scores above R104Q in both preparations (+1.652 and +1.504 against +1.427 and +1.451), a variant with two independent functional demonstrations of severe loss of function. A107D sits three residues from R104, in the same buried acidic neighborhood.

Tier 2, exceeding the threshold without introducing proline or glycine.

Variant ClinVar Rosetta OpenMM RSA All 24 frames?
I120T VCV000201430 +1.260 +1.153 43.6% yes
I102N VCV004728778 +1.015 +1.834 9.2% yes
V95L VCV000532111 +0.966 +1.207 0.0% yes
P116S VCV003619346 +1.030 +0.929 25.3% no, 22 of 24
M28T VCV002758252 +0.924 +1.048 31.0% yes
P85T VCV001518544 +1.043 +0.846 23.4% no, 19 of 24
F117L VCV001305580 +0.829 +0.855 25.4% no, 19 of 24

These rest on the exploratory positive-score selection rule rather than the specific charge-introduction signature, so I treat them as a weaker tier. V95L sits at 0.0% relative solvent accessibility, completely buried, and the substitution adds 26.7 ų of side-chain volume with nowhere in the pocket to go. P85T is a direct neighbor of D84, the buried carboxylate that coordinates R104 in the wild-type channel; in an earlier, unpublished structural analysis of my own, the P85 backbone amide appeared as one of two residual hydrogen-bond donors to that carboxylate, a plausible route by which a substitution here perturbs the same site, though I have not independently confirmed that with an additional method.

Tier 3, proline or glycine substitutions, reported but not treated as chemically interpretable.

Variant ClinVar Rosetta OpenMM RSA All 24 frames?
R104G VCV000067777 +2.018 +2.081 11.7% yes
L128P VCV000840382 +1.895 +2.351 6.8% yes
R104P VCV003691438 +1.682 +2.279 11.7% yes
T101P VCV000977688 +1.905 +1.653 17.8% yes
H118P VCV002774498 +1.402 +1.361 42.1% yes
G77E VCV004076452 +1.293 +1.285 34.0% yes
G77R VCV000919345 +1.164 +1.409 34.0% yes
R121G VCV003385973 +0.896 +1.528 11.3% yes

ThermoMPNN models side-chain packing, not backbone conformational entropy, so a large positive value for a substitution into or out of proline or glycine can be right for reasons the model does not represent. An earlier position-104 correlation analysis of my own excluded proline and glycine substitutions for the same reason. R104G and R104P sit at the same residue already addressed by functional work in the author's related studies (see competing interests), and R121G sits at the R121W position, where a dominant-negative effect is documented though contested between the two cited sources.

Fourteen of the 17 exceed the threshold in every one of the 24 relaxed frames. P116S, P85T and F117L clear it in only 22, 19 and 19 of 24 frames respectively, and I treat these three as the weakest of the set.

The clinical picture at position 104 argues for caution in reading any of this as severity ranking. R104W is ClinVar Pathogenic or Likely pathogenic with nine submitters and no conflicts, and it is not nominated (Rosetta -0.015, OpenMM +0.627), while R104G and R104P at the same residue are. Folding stability does not order clinical severity here.

A second illustration sits one residue away. Three variants of uncertain significance at D84, the buried carboxylate that accepts R104's contact in the wild-type structure, all score near zero: D84N (-0.344, -0.246), D84G (-0.067, -0.115), D84V (-0.067, +0.264). These are low predictor scores at the same structural pocket. Without independently measured endpoints for these substitutions, they establish neither benign status nor a measured sensitivity limit.

Exploratory conservation bands

A separate conservation analysis describes variation across paralogues. The historical domain set contains 172 missense records, 141 of them uncertain; the difference from the 131 used above reflects analysis-specific filtering, including structural placement on 8VYJ for the stability score.

Conservation was scored from a nine-paralogue alignment using BLOSUM62, on a 0 to 9 scale (the alignment source and version are not recorded beyond this description). The historical comparison reports separation between pathogenic- and benign-labelled positions (p = 0.027), and reports 41% of the domain at highly conserved positions. Both benign controls have low scores, R34C at 2 of 9 and V125L at 1 of 9. These observations describe the dataset; two benign controls do not calibrate a clinical evidence rule in either direction.

Conservation band VUS count Interpretation
2 of 9 or lower 41 low-conservation band; clinical meaning uncalibrated
3 to 7 of 9 41 intermediate band; clinical meaning uncalibrated
8 of 9 or higher 59 high-conservation band; clinical meaning uncalibrated

Forty-one of 141 uncertain variants sit at positions as variable as, or more variable than, either benign control.

I checked whether this filter merely restates proximity to the disordered N-terminus, since low-conservation variants cluster early in the domain. Across residues 12 to 130, residue number and conservation correlate weakly and not significantly (Spearman ρ = 0.17, p = 0.065). Conservation varies substantially within every segment of the domain (residues 12 to 45, mean 4.74 of 9; 46 to 80, mean 5.09; 81 to 115, mean 7.00; 116 to 130, mean 5.93), and the low- and high-conservation variant groups overlap in position, spanning 15 to 125 and 17 to 126 respectively. I read this as evidence against a simple positional proxy.

The 41 low-conservation variants are not thereby benign, and the 59 high-conservation variants are not thereby pathogenic. A variable position can still host a pathogenic substitution, including changes affecting backbone structure, a binding contact or splicing. The band counts are descriptive; they are not a null model or an estimate of chance discoveries. This analysis does not establish a clinical evidence rule or reclassify variants.

Classified-variant comparison and validation overlap

The historical account describes 20 already-classified variants as withheld from construction of the threshold and tiers. However, R34C, V125L, R104W and R121W also appear in the calibration-gate table above, so the named comparison is not an independent check using variants the pipeline never saw. Zero of three named benign-leaning variants were nominated (R34C -0.016, Q90K +0.129, V125L -0.258), and zero of three named pathogenic-leaning variants were nominated (R104W -0.015, R121Q +0.598, R121W -1.263). These are descriptive results, not independent estimates of generalization performance. The overlap does not establish predictor-training membership or prove that the threshold was tuned on these variants. Independent validation requires documented set membership and separation by residue position, with measured stability, trafficking and current treated as distinct endpoints.

What would change my mind

Independent functional testing could evaluate whether the nomination list enriches for experimentally abnormal variants. Calibration requires a versioned score table, independently sourced endpoint labels and position-separated validation; measured stability, trafficking and current must be distinguished. Neither the 131-variant stability selection nor the 141-variant conservation comparison supplies per-variant null probabilities or an estimated false-discovery fraction. Agreement between the two structural preparations assesses preparation sensitivity using the same predictor, not independent predictive validation.

The 114 stability non-nominations do not establish benign status. Neither the 41 low-conservation variants nor the 100 outside that band has a clinical classification established by this analysis. The lists remain exploratory pending independent validation.

Data availability

Structural data: PDB entry 8VYJ, chain A, with associated cryo-EM map EMD-43662, both from public archives. PDB entry 2OCJ was used for the independent sign-convention check. Variant data: ClinVar accessions are listed by VCV number in the tables above; the domain was queried through NCBI E-utilities, but the exact query string and retrieval date are not recorded in my working notes and I do not reconstruct them here. The predictor is ThermoMPNN (Kuhlman Lab), scored with its default bundled checkpoint. Corrected, 7 August 2026. The sentence this replaces was false, and it is corrected here rather than left to be discovered. Version 1 said "All derived tables are deposited as a single archive … They comprise the full 151-variant scoring run, the calibration gate, the tiered nomination list and the conservation filter." One of those four is deposited. Three are not, and they are not recoverable.

What is deposited, in the archive at 10.5281/zenodo.21799234 (cite version 2, 10.5281/zenodo.21840036, which is the version these filenames refer to):

What is NOT deposited, stated plainly because naming a file does not make it exist. There is no 151-row ThermoMPNN scoring table, no calibration-gate table and no tiered nomination list in the archive. They were searched for on 7 August 2026 across the whole project tree and the compute host and no copy of any of the three exists. The scoring run was not preserved as a table; its outputs survive only as the values printed in this paper. CARTDDG_CALIBRATION_FAILURE.csv is in the archive but it is a Rosetta cartesian_ddg calibration table belonging to papers 1 and 8, not this paper's gate, and it should not be mistaken for it.

Effect of this revision. Historical scores, nomination counts and reported benchmarks are retained. Their interpretation changes as described above: the threshold is exploratory, the named classified-variant comparison is overlapping, and predictive accuracy is not established. The full scoring run cannot be recomputed from the deposited material described here. The deposited conservation tables support inspection of the reported comparison, not validation of a clinical classification rule. The identifier is recorded in DATA_DOI.txt alongside this manuscript.

Archive versioning. The concept DOI 10.5281/zenodo.21799233 always resolves to the current version of the data archive and is the identifier to follow for access. The version current at the time of this revision is version 2, 10.5281/zenodo.21840036. Version DOIs cited elsewhere in this manuscript name the specific version read and are deliberately not rewritten.

Competing interests

I disclose that I am a heterozygous carrier of the SCN5A variant p.Arg104Gln (R104Q), which appears as a reference point in the scoring above.

Use of AI tools

This work was carried out with AI coding and research assistants (Anthropic Claude, via Claude Code). That use is disclosed here rather than left to inference.

Analysis code. The great majority of the analysis code in this project -- parsers, genome scans, regeneration scripts and verification scripts -- was written by an AI assistant working to my specification. I set what each script had to compute, chose the thresholds and the decision rules, and checked the output against the claims it is used to support.

Manuscript text. The prose of this manuscript was drafted by an AI assistant. I directed the drafting and revised the result, and I am responsible for every claim it makes.

Scientific decisions. The questions asked, the thresholds set, what was allowed to count as a refutation, and what was published were mine.

Verification, which does not depend on any of the above. Where a claim in this manuscript is regenerable from deposited inputs, the script that regenerates it and that script's own output are in the data deposit. Reproduction does not require trusting any account of who wrote what.

What no AI system did. No AI system generated, altered or selected any experimental measurement; this project contains no wet-lab data of any kind. All primary literature cited was retrieved from PubMed, PMC and publisher sources. Every reference in this manuscript has been machine-resolved against its own record, including a check that each PMID's first author and year match the author and year printed beside it in the text.

References

  1. Moreau A, et al. 2012. PMID 22529811.
  2. Clatot J, et al. 2012. PMID 22739120.
  3. Wang Y, et al. 2020. PMID 32815768.
  4. Levy-Nissenbaum E, Eldar M, Wang Q, et al. Genetic analysis of Brugada syndrome in Israel: two novel mutations and possible genetic heterogeneity. Genet Test. 2001;5(4):331-334. PMID 11960580. DOI 10.1089/109065701753617480. (Resolved 10 August 2026.)
  5. Tien MZ, Meyer AG, Sydykova DK, Spielman SJ, Wilke CO. Maximum allowed solvent accessibilites of residues in proteins. PLoS One. 2013;8(11):e80635. PMID 24278298. DOI 10.1371/journal.pone.0080635. (Resolved 10 August 2026.)
  6. Protein Data Bank entry 8VYJ, chain A.
  7. Electron Microscopy Data Bank entry EMD-43662.
  8. Protein Data Bank entry 2OCJ, used for the independent sign-convention check.
  9. Dieckhaus H, Brocidiacono M, Randolph NZ, Kuhlman B. Transfer learning to leverage larger datasets for improved prediction of protein stability changes. Proc Natl Acad Sci U S A. 2024;121(6):e2314853121. PMID 38285937. DOI 10.1073/pnas.2314853121. (ThermoMPNN; the tool's bundled default checkpoint carries no separate version label. Publication resolved 10 August 2026.)