Version of record: 10.5281/zenodo.21799870, published 5 August 2026. That identifier is the citable address for this paper and it resolves at https://doi.org/10.5281/zenodo.21799870. It is a version identifier; Zenodo minted a second one that resolves to all versions, and the version identifier is the one to cite.
What this file is. This project's authoritative copy of the manuscript,
SUBMIT_THESE/papers/PUBLISH_10_PENETRANCE_COVARIATES.md, which is the file the deposited PDF was built from. Synced 7 August 2026 byscripts/sync-manuscripts.mjs, which copies the source byte for byte and prepends this note. Nothing in the manuscript below has been rewritten for the website.Warning, and it points at the published record rather than at this page. This is the most serious defect found in the set, because it is a statement about the world rather than a wrong number. The data availability statement in the version deposited on 5 August 2026 told readers that this note's derived tables were deposited in the project's data archive,
10.5281/zenodo.21799234, and named four of them: the pairwise geometry table, the per-residue neighbour count table, the per-residue sigmoid weight table, and the recompute output. None of the four were in that archive. Of the ten papers that archive supported, this note was the only one with no file in it of any kind. A reader who downloaded the archive to check these numbers found nothing to check, and nothing to explain the absence. A wrong number can be caught by recomputing it; a false data availability statement can only be caught by downloading the archive and finding nothing there, and having found nothing, the reader cannot tell whether the archive is wrong, the paper is wrong, or they are looking in the wrong place.It was not a filename error and it could not be fixed by locating the files. The code that produced the originals is not in the project's file tree and is not on its compute host. Both were searched on 6 August 2026 and returned only third-party library files: there is no
.Rfile and no R installation on either machine. This note's own Methods already record that the software used "is not recorded in the source material".Three of the four tables were regenerated and the fourth is declared missing. The pairwise geometry, per-residue neighbour count and per-residue sigmoid weight tables are pure geometry over a public structure, PDB 8VYJ chain A, and the method is fully specified in this note. They were regenerated on 6 August 2026 by
p10_regen_geometry.py, written from the note's stated method, and thirty-five of the thirty-seven quantities printed in this note reproduce exactly from them: every distance, every AUC, every correlation and every per-residue count named in the text. One deviation is recorded rather than silently adjusted: side-chain-centroid precision at 8 Å reproduces as 0.606 against the 0.608 printed, a single borderline pair. The fourth table, the recompute output, cannot be produced and is declared missing rather than promised, because it requires the kroncke-labBayes_BrS1_PenetranceR pipeline and a refit of its expectation-maximization step. The regenerated tables and the script are staged for version 2 of the archive and have not been uploaded, so at the moment of writing the archive at that identifier still holds nothing for this note.Two defects in the note itself were found by that reproduction, and both cut against its own argument. First, the resolved N-terminal domain contains four salt bridges, not two: under the definition this note states in its own table, acidic oxygen to basic nitrogen below 4 Å, it holds R14-E78 at 3.18 Å, E25-K91 at 3.01 Å, E30-R34 at 2.79 Å and D84-R104 at 3.79 Å. The two that were not counted are both captured by the 8 Å centroid cutoff, so the finding that D84-R104 is missed stands and the denominator does not. The miss-rate table's salt-bridge row moves from 2 / 1 / 50% to 4 / 1 / 25%, and the abstract with it. Second, the per-residue neighbour counts include cross-domain contacts and the deposited version did not say so; counting only within the domain gives Pearson r = 0.756 and Spearman ρ = 0.779 against the 0.778 and 0.791 printed, so none of those figures is reproducible from the method as the deposited version stated it. This copy states the rule.
No conclusion in this note changes. The centroid covariate still misses one close contact in eleven, the misses still concentrate in long-range side-chain contacts, R104-D84 is still among them, and substituting the real 8VYJ geometry still moves the R104Q penetrance estimate by only 0.9 points. The null result is untouched.
If a figure on this page disagrees with the same figure at the identifier above, this page is the corrected one. The full divergence, and what a version-2 deposit would have to include, is recorded in
SUBMIT_THESE/ZENODO_DIVERGENCE_20260806.md. (This paragraph used to open by asserting that no version 2 had been deposited and nothing had been uploaded. That was true when written on 6 August 2026 and is not a claim a generated page can keep true, because it would turn false the moment anything is deposited and nothing here would notice. The sentence is removed rather than updated: to find out what is deposited, resolve the identifier, which is the only source that cannot go stale.)None of this is peer reviewed, and none of it has been through a wet lab. No cell has been edited and no current has been recorded for this variant by this project. Every therapeutic statement in the manuscript below is a prediction.
Residue centroid distance misses close atom contacts in the SCN5A N-terminal domain, and real cryo-EM geometry does not move the R104Q penetrance estimate
Ethan Bradley
Independent researcher, no institutional affiliation
ORCID: 0009-0008-8925-7975
Abstract
Residue centroid distance, the structural covariate used in a published Brugada syndrome penetrance model, misses close atom contacts within the Nav1.5 (SCN5A) N-terminal domain, and the misses concentrate in long side chain and salt bridge contacts rather than spreading evenly across the domain. Using PDB 8VYJ chain A (cryo-EM, approximately 3.6 Å, map EMD-43662, residues 12 to 130), I computed closest heavy atom and centroid distances for all 5,672 residue pairs at sequence separation of two or more. Of 218 pairs in close atom contact (closest heavy atom below 4.0 Å), 20, or 9.2 percent, fall beyond an 8 Å centroid cutoff and are invisible to that covariate; sensitivity is 0.908 and precision 0.662. The domain contains four salt bridges, and one of them, R104 to D84 (3.79 Å between closest heavy atoms, 9.22 Å between centroids), is among the missed contacts. With only four present, that is one miss out of four and not a rate. Supplying the real 8VYJ geometry for residue 104 in place of the model's sequence-distance fallback, and refitting the full expectation-maximization step, changes the R104Q penetrance estimate from 42.64 percent to 41.70 percent, a shift of 0.9 points that holds across neighbor cutoffs from 12 to 25 Å. That is a null result: the salt bridge carries little weight once it passes through a centroid-based covariate, so correcting the geometry for one residue does not move the estimate.
A key to the terms used here
- SCN5A is the gene for the heart's main sodium channel; Nav1.5 is the protein. R104Q is arginine at position 104 replaced by glutamine, the variant used as the worked example.
- Penetrance is the proportion of people carrying a variant who actually develop the condition. A penetrance model estimates it from a variant's properties, and covariates are the properties fed in.
- A residue is one amino acid in a protein chain. Its centroid is the average position of its atoms, a single point standing in for the whole side chain.
- Centroid distance between two residues is the distance between those two summary points. The finding here is that this summary hides close contacts, because two residues can be far apart on average while individual atoms nearly touch. Averaging a shape to a point loses the parts that stick out.
- Structural density, or the neighbour count, is how many other residues sit near a given one. It is the covariate this paper tests.
- A salt bridge is an attraction between a positively and a negatively charged side chain. These are formed at the tips of long side chains, which is why a centroid summary misses them disproportionately.
- An ångström, Å, is a ten-billionth of a metre; atoms in contact sit a few ångströms apart.
- Cryo-EM is a microscopy method producing a 3D model of a protein. PDB 8VYJ is the model used here and EMD-43662 its density map. Real measured geometry replaces the idealised geometry the published model assumed.
- Sigmoid weighting is a way of counting neighbours that fades smoothly with distance instead of applying a hard cut-off.
- Brugada syndrome is an inherited arrhythmia condition linked to reduced cardiac sodium current.
Why the definition of "structural neighbor" matters here
Structure-based penetrance models for arrhythmia genes use a structural-density covariate: a distance-weighted count of neighboring residues, sitting alongside sequence-conservation covariates in the same fit. The distance in that covariate is conventionally measured between residue centroids, the mean position of a residue's atoms. For a long side chain that makes contact through its tip rather than its base, the centroid sits behind the point of contact, and a covariate built on that distance can average away the interaction it is meant to register. R104 in the Nav1.5 N-terminal domain forms a salt bridge to D84. R104Q is my own variant, and it is the case used to test whether this concern is real or academic. The question below is answered in two parts: does the centroid definition actually miss contacts of this kind in this domain, and does the answer change anything once it is fed back into the published model.
Methods
Structure. PDB 8VYJ, chain A, a cryo-EM reconstruction at approximately 3.6 Å resolution, deposited with electron microscopy map EMD-43662 (Biswas et al., 2025). Residues 12 to 130 are modelled, 108 residues in total; residues 38 to 48 are unresolved and excluded.
Pairwise geometry. For every pair among the 108 modelled residues with sequence separation of two or more (5,672 pairs), I computed five distances from heavy atoms only, excluding hydrogens and alternate conformers: closest heavy atom distance, all-atom centroid distance, side-chain centroid distance, Cα to Cα distance, and Cβ to Cβ distance (Cα substituted for Cβ at glycine). A pair was called a close atom contact when the closest heavy atom distance was below 4.0 Å. The centroid cutoff tested against that definition was 8 Å. The specific software used to parse coordinates and compute these distances is not recorded in my working notes beyond that it operates directly on the PDB coordinate file; no package name or version is given, and I report that gap rather than guess at one.
Validation. Computed distances for R104 to D84 and R104 to D82 were checked against independently recorded values for the same two pairs (3.79 Å closest atom and 9.22 Å centroid for R104-D84; 4.78 Å and 8.07 Å for R104-D82) and agreed to 0.01 Å.
Penetrance model. The BrS1 penetrance model and its code are distributed in the repository kroncke-lab/Bayes_BrS1_Penetrance on GitHub. I ran the pipeline as published, using its own files func_dist_seq.R, distance_file, and BrS1_data.RData, following its steps in order: weighted penetrance, a method-of-moments Beta prior, a per-variant posterior, the funcdist structural-density term, an expectation-maximization regression on the six-feature set eaRate, blastpssm, provean_score, pph2_prob, ipeak, and feat_dist_w, a variance scaled by 20, and the resulting posterior. No repository access date or software version is recorded in my working notes, so neither is stated here as fact. The pipeline's own EM is capped at 10 iterations and plateaus at a δ of approximately 0.63; that behavior was reproduced, not altered.
Sequence fallback. Where the pipeline's own distance file lacks resolved structure for a residue, because it is built from transmembrane-domain templates that do not include the N-terminal domain, it substitutes a sequence-distance approximation of 3.8 times the square root of the residue separation. Residue 104 is one such residue.
Recompute. I calculated all-atom centroid distances from residue 104 to every chain A residue within 25 Å in 8VYJ. To confirm which distance convention the pipeline itself uses, I compared Cα-Cα, Cβ-Cβ, side-chain centroid, and all-atom centroid distances against the pipeline's own distance file on three pairs it does resolve (120-126, 120-117, 120-178). All-atom centroid matched best, with a mean absolute difference of 0.5 Å. I then replaced the pipeline's entry for residue 104 with the real 8VYJ neighbor list, including a self-distance of 0 Å to match the pipeline's own convention, so that R104W and R104G continue to count residue 104 as a neighbor of itself at weight 0.5, identically to the sequence fallback, and re-ran the full EM.
Centroid distance misses one in eleven close atom contacts, and the misses are not random
Of the 5,672 pairs evaluated, 218 form a close atom contact (3.84%). Of those 218, 20 (9.2%) have a centroid distance above the 8 Å cutoff and are therefore invisible to a centroid-based covariate at that threshold. The most extreme case is K62 to K100: 3.14 Å between the closest heavy atoms, 9.64 Å between centroids. Treating the 8 Å centroid cutoff as a classifier for a real contact gives a sensitivity of 0.908 and a precision of 0.662: the covariate misses roughly one contact in eleven, and about a third of what it does count as a neighbor is not actually in contact.
The miss rate is not uniform. It rises with sequence separation and with the involvement of side chains:
| Contact class | n | missed | miss rate |
|---|---|---|---|
| sequence separation 2-4 | 132 | 7 | 5.3% |
| sequence separation 5-11 | 29 | 3 | 10.3% |
| sequence separation ≥12 | 57 | 10 | 17.5% |
| ≥12 and side-chain to side-chain | 26 | 6 | 23.1% |
| salt bridges (acidic O to basic N, <4 Å) | 4 | 1 | 25% |
The salt bridge row is built on four observations, not a sample. There are four salt bridges in the resolved N-terminal domain, and one of them is missed. That is one miss out of four, and it should be read as exactly that rather than as a 25% rate that would generalize to other salt bridges. The bridge that is missed is R104 to D84: 3.79 Å between the guanidinium nitrogen and the carboxylate oxygen, 9.22 Å between residue centroids. The other three are all captured, and comfortably: E25 to K91 at a 6.74 Å centroid distance, R14 to E78 at 6.13 Å, and E30 to R34 at 4.29 Å. (This row read "2 / 1 / 50%" until 6 August 2026 and was wrong; see the correction section below. One miss out of four is a weaker claim than one out of two, and the direction of the error is against this paper's own argument rather than for it.)
Six long-range side-chain contacts are missed in total: L21-F117, R27-F86, L67-L83, D84-R104, F93-R121, F105-L128. Two of the positions involved, 104 and 121, carry variants described elsewhere as dominant-negative for channel function; this note does not add a new citation for that claim beyond noting the coincidence of position.
Cross-domain contacts, meaning contacts between an N-terminal-domain residue and the rest of the channel, follow the same pattern. R104 has a relative solvent accessibility of 11.7% in the full channel but 32.6% when the N-terminal domain is considered in isolation, consistent with burial against the channel body. Extending the pairwise comparison to N-terminal-domain-versus-rest-of-chain pairs gives 21 close atom contacts across 11 domain residues, of which 2 (9.5%) are missed by the centroid measure, close to the within-domain rate. R104's own cross-domain contact to R179 (3.65 Å atoms, 8.78 Å centroids) is missed; its contact to F186 (3.31 Å, 6.27 Å) is captured.
Per-residue neighbor counts disagree even where the pairwise picture looks stable
A model does not consume individual pair distances. It consumes a per-residue neighbor count, and that count diverges more than the pairwise miss rate suggests. These per-residue counts include each domain residue's contacts to the rest of chain A as well as its contacts within the domain; the pairwise table above is domain-internal only. (That distinction was not stated in the version published 5 August 2026 and is added here on 6 August 2026, because without it none of the correlation figures in this paragraph can be reproduced. Domain-internal counts alone give r = 0.756 and ρ = 0.779, not the values below.) Recomputed both ways across the domain, contact-based and centroid-based per-residue counts correlate at Pearson r = 0.778 and Spearman ρ = 0.791, with a mean absolute difference of 1.81 neighbors and a maximum difference of 6. Ninety-one of 108 residues, 84%, receive a different neighbor count depending on which definition is used. R104 itself has 6 contact-based neighbors against 8 centroid-based neighbors. The largest disagreements sit elsewhere in the domain: G77 (2 versus 8), I94 (4 versus 9), S106 (5 versus 10), P79 (3 versus 8), and L96 (6 versus 11), all in loop and strand regions where several of the domain's characterized variants are found.
Swapping in a different single-point summary does not fix this
The obvious fix is to move the summary point, for instance to a side-chain centroid instead of an all-atom centroid. Benchmarked against the same close-atom-contact target, that trade is not favorable:
| Measure | AUC | Sensitivity at 8 Å | Precision at 8 Å |
|---|---|---|---|
| all-atom centroid | 0.993 | 0.908 | 0.662 |
| side-chain centroid | 0.984 | 0.697 | 0.608 |
| Cα-Cα | 0.988 | 0.849 | 0.631 |
| Cβ-Cβ | 0.990 | 0.881 | 0.667 |
Side-chain centroid recovers 9 of the 20 missed contacts, including R104-D84 (9.22 Å falling to 6.76 Å), but it loses ground elsewhere and ends up the weakest of the four measures overall. All-atom centroid remains the best average proxy of the four tested. The problem here is not that the wrong point was chosen to summarize a residue. It is that any single-point summary discards information that only a closest-atom measurement retains, namely whether two residues actually touch.
Sigmoid weighting softens the miss without closing it
Published models weight neighbors with a smooth distance decay rather than a hard cutoff, so the practical failure is one of weight rather than binary exclusion. Applying a generic logistic weight (midpoint 7 Å, scale 1 Å) to both distance definitions shows the same pattern in continuous form: centroid distance gives a real sub-4 Å contact a mean weight of 0.598 (minimum 0.067), while contact distance gives the same contacts a mean weight of 0.972 (minimum 0.953). For R104-D84 specifically, centroid distance assigns a weight of 0.098; contact distance assigns 0.961, roughly ten times more. As a share of R104's total structural weight, the salt bridge contributes about 1.9% under the centroid definition by this study's own sigmoid computation. Reading the published model's own weights for the same residue puts the same bond at roughly 2.3%, an independent route arriving at a similar order of magnitude. Across the whole domain, the two per-residue covariates correlate at ρ = 0.868, close enough that aggregate model fit could plausibly survive the switch, and far enough apart that individual buried-contact residues are consistently under-weighted.
Substituting real 8VYJ geometry for residue 104 does not move the R104Q estimate
Reproducing the published pipeline exactly, before changing anything, returns R104Q's shipped values to the digit: structural density feat_dist_w = 0.21811, α_g = 5.233, β_g = 13.767, and a BrS1 penetrance of 0.4264 (90% credible interval 0.267 to 0.593).
Replacing the sequence-distance fallback for residue 104 with the real 8VYJ neighbor geometry, and re-running the full EM, gives:
| structural density (feat_dist_w) | BrS1 penetrance | 90% credible interval | |
|---|---|---|---|
| Baseline (sequence fallback) | 0.2181 | 0.4264 | 0.267-0.593 |
| 8VYJ (real geometry) | 0.2051 | 0.4170 | 0.259-0.583 |
| Difference | -0.013 | -0.009 |
R104Q's penetrance estimate moves from 42.64% to 41.70%, a shift of 0.9 points. That result holds at every neighbor cutoff tested, 12, 15, 20, and 25 Å, returning 0.417 in each case. The same-residue variants shift in the same small, downward direction: R104W moves from 0.437 to 0.409, and R104G from 0.490 to 0.431.
This is a null, and the salt bridge analysis above explains why. D84 sits at 9.22 Å by residue centroid, so under the covariate definition the model actually uses, the bridge contributes on the order of 0.010 of R104's roughly 0.43 total structural weight, about 2.3%, matching the independent 1.9% to 2.3% estimate from the sigmoid weighting above. A structural detail about one buried contact does not have much room to move a covariate that averages it away in the first place. The correct reading is that the covariate did not transmit the geometry, not that the geometry is unimportant.
What would overturn this
A few specific observations would change the conclusions here, and it is worth naming them rather than leaving the claim unfalsifiable. If a larger set of salt bridges, drawn from other SCN5A domains or other resolved structures, showed that most salt bridges are in fact captured by an 8 Å centroid cutoff, the one-of-four result here would be revealed as noise from a domain that happens to have only four such bridges, rather than a systematic property of the covariate. Three of the four in this domain already are captured, so that overturning is nearer than the version of this paragraph published on 5 August 2026 implied. If refitting the full EM with closest-heavy-atom distances in place of centroid distances across the whole domain, not just residue 104, left the per-residue influence pattern unchanged, the covariate-definition explanation offered for the null would be wrong and some other factor would need to be found. If removing the pipeline's built-in 10-iteration cap produced a materially different R104Q estimate under the 8VYJ geometry, the stability reported above would not hold as stated. Finally, this analysis uses a single static cryo-EM conformer; if the R104-D84 bridge is not maintained across the channel's accessible conformational range, a comparison across multiple structures or models could change the geometric picture the covariate is being asked to summarize. No such comparison was attempted here.
Two further limits apply to the benchmark itself. The 4.0 Å and 8 Å thresholds are conventional, not derived, and while the qualitative pattern (misses concentrating in long-range side-chain contacts) is not sensitive to small changes in either threshold, the exact counts reported here are. And the side-chain-centroid comparison in the table above uses the same 8 Å cutoff as the all-atom centroid measure for comparability, which is not necessarily that measure's own best-performing cutoff.
Correction, 6 August 2026
The data availability statement in the version of this note deposited at
10.5281/zenodo.21799871 on 5 August 2026 was false. It told readers that this note's derived tables
were in the project's data archive (10.5281/zenodo.21799234) and named four of them: the pairwise
geometry table, the per-residue neighbor count table, the per-residue sigmoid weight table, and the
recompute output. None of the four were in that archive. Of the ten papers the archive supported, this
one was the only one with no file in it at all. Anyone who downloaded version 1 of the deposit to check
the numbers in this note found nothing to check, and nothing to tell them why.
It was not a filename error and it could not be fixed by locating the files. The code that produced
the original tables is not in the project's file tree and is not on its compute host; both were searched
on 6 August 2026 for geometry, neighbour, sigmoid and recompute artefacts and returned only third-party
library files. There is no .R file and no R installation on either machine. This note's own Methods
already record that the software used "is not recorded in my working notes", and that gap turned out to
be larger than it read.
What was done about it. Three of the four named tables are geometry over a public structure and were
regenerated on 6 August 2026 by a script written from the method stated above, deposited with that script
and with its verification output in version 2 of the data archive, together with a fourth table of
cross-domain contacts that this note's per-residue counts require and did not name. Thirty-five of the
thirty-seven quantities printed in this note reproduce exactly from the regenerated tables, including
every distance, every AUC, every correlation and every per-residue count named in the text. The fourth
named table, the recompute output, is still not deposited and is not reconstructible: it needs the
kroncke-lab/Bayes_BrS1_Penetrance pipeline and a refit of its expectation-maximization step, and no copy of
that pipeline and no R installation exists in this project. The files added are
P10_PAIRWISE_GEOMETRY_8VYJ.csv, P10_PER_RESIDUE_NEIGHBOR_COUNTS.csv,
P10_PER_RESIDUE_SIGMOID_WEIGHTS.csv and P10_CROSS_DOMAIN_CONTACTS.csv, with the script, its
verification output and a provenance note beside them. The structural density, α_g, β_g and penetrance
values in this note exist only as printed here. The data availability statement below says so.
Two defects in the note itself were found by that reproduction, and both are corrected above.
| Where | Was | Is | Why |
|---|---|---|---|
| Abstract, and the salt-bridge row of the miss-rate table | "exactly two salt bridges", 2 / 1 / 50% | four salt bridges, 4 / 1 / 25% | Under the definition this note states in its own table — acidic oxygen to basic nitrogen below 4 Å — the resolved domain holds four: R14-E78 (3.18 Å), E25-K91 (3.01 Å), E30-R34 (2.79 Å) and D84-R104 (3.79 Å). The two that were not counted are both captured by the 8 Å centroid cutoff |
| Per-residue neighbor counts | method not stated | states that the counts include cross-domain contacts | Domain-internal counts alone give r = 0.756, ρ = 0.779, mean absolute difference 1.65 and 85 of 108 residues differing, not the 0.778, 0.791, 1.81 and 91 printed. Adding cross-domain contacts reproduces all four exactly, and R104's 6-versus-8 and S106's 5-versus-10 with them |
Neither changes a conclusion, and the first weakens this note's own argument rather than strengthening it. The missed bridge is still missed, it is still the one that matters for R104, and the null result is untouched — but a miss rate of one in four is a smaller signal than one in two, and the sentence warning against reading that row as a rate now has to carry less weight, not more. One further deviation is recorded and is trivial: side-chain-centroid precision at 8 Å reproduces as 0.606 against the 0.608 printed here, one borderline pair, with three non-contact pairs sitting between 8.00 and 8.02 Å by that measure.
None of this section is in the record deposited on 5 August 2026, which carries the false data availability statement, the two-salt-bridge count and the unstated per-residue method.
Data availability
This statement was rewritten on 6 August 2026 because the version published on 5 August 2026 named four deposited tables of which none were in the archive; see the correction section above. The structure used is PDB 8VYJ, chain A, with electron microscopy map EMD-43662, both public accessions. The penetrance model and its code are the public repository kroncke-lab/Bayes_BrS1_Penetrance on GitHub, specifically the files func_dist_seq.R, distance_file, and BrS1_data.RData; no access date is recorded for that repository in my working notes. The derived tables are deposited as a single archive with a permanent identifier, recorded in DATA_DOI.txt alongside this manuscript and to be cited as the data source, and from version 2 of that archive onward they comprise seven files: P10_PAIRWISE_GEOMETRY_8VYJ.csv, the pairwise geometry table, all 5,672 residue pairs at sequence separation two or more with closest-heavy-atom, all-atom centroid, side-chain centroid, Cα-Cα and Cβ-Cβ distances, the contact and neighbour flags, both sigmoid weights and a salt-bridge flag carrying its acidic-oxygen-to-basic-nitrogen distance; P10_PER_RESIDUE_NEIGHBOR_COUNTS.csv, the per-residue neighbour count table over 108 residues with within-domain, cross-domain and total columns kept separate; P10_PER_RESIDUE_SIGMOID_WEIGHTS.csv, the per-residue sigmoid weight table over the same residues at logistic midpoint 7 Å and scale 1 Å; P10_CROSS_DOMAIN_CONTACTS.csv, the 21 close atom contacts between the domain and the rest of chain A, which the per-residue counts require and which the published version did not name; p10_regen_geometry.py, the script that produces all four from the public 8VYJ coordinates; P10_VERIFICATION_OUTPUT.txt, that script's own print-out setting each reproduced quantity beside the value printed here; and P10_REGENERATION_NOTE.md, which states that these tables are a 6 August 2026 reconstruction rather than the 4 August originals and why the originals are unrecoverable. The recompute output is not deposited and is not reconstructible: it would hold the structural density and penetrance values under the baseline and 8VYJ geometries, and producing it requires the kroncke-lab pipeline and a refit of its expectation-maximization step, which this project can no longer run, so those values are available only as printed in this note and should be treated accordingly.
Archive versioning. The concept DOI 10.5281/zenodo.21799233 always resolves to the current version
of the data archive and is the identifier to follow for access. The version current at the time of this
revision is version 2, 10.5281/zenodo.21840036. Version DOIs cited elsewhere in this manuscript name the
specific version read and are deliberately not rewritten.
Competing interests
I am a heterozygous carrier of SCN5A R104Q (NM_000335.5:c.311G>A, p.Arg104Gln), the variant used as the worked example throughout this note.
Use of AI tools
This work was carried out with AI coding and research assistants (Anthropic Claude, via Claude Code). That use is disclosed here rather than left to inference.
Analysis code. The great majority of the analysis code in this project -- parsers, genome scans, regeneration scripts and verification scripts -- was written by an AI assistant working to my specification. I set what each script had to compute, chose the thresholds and the decision rules, and checked the output against the claims it is used to support.
Manuscript text. The prose of this manuscript was drafted by an AI assistant. I directed the drafting and revised the result, and I am responsible for every claim it makes.
Scientific decisions. The questions asked, the thresholds set, what was allowed to count as a refutation, and what was published were mine.
Verification, which does not depend on any of the above. Where a claim in this manuscript is regenerable from deposited inputs, the script that regenerates it and that script's own output are in the data deposit. Reproduction does not require trusting any account of who wrote what.
What no AI system did. No AI system generated, altered or selected any experimental measurement; this project contains no wet-lab data of any kind. All primary literature cited was retrieved from PubMed, PMC and publisher sources. Every reference in this manuscript has been machine-resolved against its own record, including a check that each PMID's first author and year match the author and year printed beside it in the text.
References
- Biswas R, Lopez-Serrano AL, Purohit A, et al. Structural basis of human Na(v)1.5 gating mechanisms. Proc Natl Acad Sci U S A. 2025;122(20):e2416181122. PMID 40366698. DOI 10.1073/pnas.2416181122. Structure deposited as PDB 8VYJ, chain A (deposited 8 February 2024, released 12 February 2025); electron microscopy map EMD-43662. (Resolved 10 August 2026 from the RCSB entry's primary citation.)
- Kroncke lab. Bayes_BrS1_Penetrance. GitHub repository: kroncke-lab/Bayes_BrS1_Penetrance. Files used: func_dist_seq.R, distance_file, BrS1_data.RData. Repository access date not recorded in my working notes.