Lilienfeld 1925 Establish Reproducibility Across Specimens
# Establish Reproducibility Across Specimens: One Passing Device Is an Anecdote, Not a Result
Step 23 produced a pass, hold, or reject verdict for a single, specific specimen. Step 24 asks whether that verdict generalizes: build a population of independently fabricated devices through the identical Steps 1–14 construction sequence, run each one through the identical Steps 15–23 verification arc, and report the distribution of outcomes rather than the outcome of the one device that happened to be measured first. This is not a formality. The historical record around hand-built compound-film control structures from this era is one of isolated favorable observations that could not be repeated on demand, and a verification program that stops at a single passing specimen reproduces that exact historical pattern instead of correcting it.
A single device, however carefully measured, cannot distinguish a real, repeatable field effect from a rare favorable coincidence of geometry, film defect, and contact condition that happened to survive Steps 15–18 and produce a Step 23 pass. Only a population can make that distinction, because only a population has a rate.
## 1. Define the population and the independence it requires before building it
A population means specimens fabricated as separate instances of the full Steps 1–14 sequence, not multiple measurements of one device, and not multiple devices cut from a single sulfurization batch sharing one film-forming event. Record, for every specimen, which fabrication sub-steps were shared (a common foil lot, a common sulfurization furnace run, a common coating session) and which were independent, because a defect common to a shared sub-step can make several specimens fail or pass together for a reason that has nothing to do with the field-effect mechanism being tested.
State the population size before any specimen is measured, using a number large enough that the statistics in §4 can discriminate a real effect rate from a chance rate at a stated confidence, not a number chosen after seeing how many early specimens passed. Revising the target size upward only after an unfavorable early run is the same protocol-adaptation failure Step 21 was written to prevent, applied now at the population level instead of the single-device level.
## 2. Run every specimen through the full, unmodified Steps 15–23 arc
Each specimen receives the identical geometry inspection, film characterization, isolation check, baseline measurement, leakage measurement, frozen bias protocol, modulation test, and power-gain quantification already defined in Steps 15 through 23, with no shortcuts taken for specimens that look unpromising early and no extra scrutiny applied only to specimens that look promising early. Measuring a favorable-looking specimen more carefully than an unfavorable-looking one biases the population toward favorable verdicts before the comparison in §4 ever runs.
Carry forward every intermediate disposition, not just the Step 23 verdict. A specimen rejected at Step 18 for a conductive bridge, or at Step 20 for leakage, is a data point about the fabrication process, not a specimen to be quietly discarded and replaced without being counted. Report the full funnel: how many specimens entered Step 15, how many survived each intermediate reject point, and how many reached a Step 23 verdict at all.
## 3. Hold the frozen protocol and thresholds fixed across every specimen
The Step 21 bias list, interleaving schedule, and the $\delta I_{min}$ and $f_{leak,max}$ thresholds from Steps 19 and 20 must be the same written protocol applied to every specimen, not re-derived per specimen from that specimen's own early data. A per-specimen threshold that is quietly loosened for a device that would otherwise fail, or tightened for one that would otherwise pass by a wide margin, destroys the comparability that makes a population meaningful. If a specimen's own baseline repeatability genuinely differs enough to need a different $\delta I_{min}$, document that as a property of the specimen and carry the original population-level threshold forward as well, so both can be reported.
## 4. Report a rate with an interval, not a count of successes
Convert the per-specimen verdicts into a population-level statistic. For $n$ specimens reaching a final Step 23 disposition, with $k$ passing, report the pass rate
together with a binomial confidence interval appropriate for small $n$ — a Wilson or Clopper–Pearson interval rather than a normal approximation, since early-stage reproducibility runs are typically small. Where §1 identified specimens sharing a sub-step, do not treat them as $n$ fully independent trials; either exclude all but one specimen per shared cluster from the rate calculation, or use a clustering-adjusted estimator, and report which choice was made and why.
State plainly what the interval does and does not support. A high $\hat p$ with a wide interval at small $n$ supports cautious optimism and a larger follow-up population, not a settled claim. A low $\hat p$, or an interval that comfortably includes rates consistent with chance alignment of the confounds Step 23 already named, supports a conclusion that the Step 23 pass on the original specimen was not representative.
## 5. Define the chance-rate reference before comparing against it
The dashed reference in §4 is not an arbitrary benchmark; it must be derived from the same confounds Steps 18 through 23 were built to exclude. Estimate, from the fixture-blank and insulating-blank data already collected across Steps 18 and 20, how often a specimen with no real control-electrode effect would be expected to produce a Step 19–23 verdict pattern that looks like a pass purely from leakage variability, baseline noise, or an unflagged correlated artifact. This reference rate — not zero, and not an assumed figure — is the rate a real effect must clear.
## 6. Characterize the failing specimens, not only the passing ones
Fold the geometry, coverage, and material maps from Steps 15 through 17 back into the population analysis. If specimens that pass Step 23 share a measurable trait — a narrower minimum foil clearance, a particular film-thickness range, a specific coverage pattern — that correlation is itself a finding, and it may point toward the physical condition the field effect actually depends on, or toward a specific fabrication defect that happens to produce a pass through a path §5's reference rate does not anticipate. Report this correlation explicitly rather than treating passing specimens as an undifferentiated group.
## 7. Minimum Step 24 record
Retain:
- the pre-declared population size, its independence structure across shared fabrication sub-steps, and the written justification for both, fixed before the first specimen reached Step 15;
- the complete Steps 1–14 fabrication record for every specimen, with shared sub-steps flagged;
- the complete Steps 15–23 record for every specimen, including every intermediate reject disposition, not only specimens reaching a final Step 23 verdict;
- confirmation that the Step 19–21 thresholds and frozen protocol were applied identically across all specimens, with any documented exception stated;
- the full disposition funnel from entry to final verdict;
- the pass rate $\hat p$, its confidence interval and the method used, and the independence-adjustment method where shared sub-steps were present;
- the chance-rate reference derived from Step 18 and Step 20 blank data, and the comparison against it;
- any measured geometry, coverage, or material correlation distinguishing passing from failing specimens;
- the final reproducibility disposition for the population as a whole.
The historical geometry and claimed effect are grounded in [J. E. Lilienfeld, US Patent 1,745,175](https://patents.google.com/patent/US1745175A/en). The requirement to establish a rate across independent trials before accepting a device-level claim, rather than generalizing from a single specimen, follows the same discipline that separates a reproducible instrumental result from an isolated favorable observation throughout the low-level measurement practice referenced in the [Keithley Low Level Measurements Handbook](https://www.tek.com/en/documents/product-article/keithley-low-level-measurements-handbook---8th-edition).
## Establish Reproducibility Across Specimens’ Place in the Process Lineage
Steps 15 through 23 built a rigorous verification arc and applied it to one specimen. Step 24 now runs that same arc, unmodified, across a pre-declared population of independently fabricated specimens, tracks every intermediate disposition rather than discarding unfavorable ones, and reports a pass rate with a confidence interval against a chance-rate reference derived from the project's own blank measurements. Only a population-level result that clears that reference, with documented independence and a characterized relationship between passing specimens and their measured geometry, coverage, and material properties, extends the Step 23 amplification verdict from one device to a claim about the fabrication process as a whole.