Lilienfeld 1925 Establish Reproducibility Across Specimens

# Establish Reproducibility Across Specimens: One Passing Device Is an Anecdote, Not a Result

Step 23 produced a pass, hold, or reject verdict for a single, specific specimen. Step 24 asks whether that verdict generalizes: build a population of independently fabricated devices through the identical Steps 1–14 construction sequence, run each one through the identical Steps 15–23 verification arc, and report the distribution of outcomes rather than the outcome of the one device that happened to be measured first. This is not a formality. The historical record around hand-built compound-film control structures from this era is one of isolated favorable observations that could not be repeated on demand, and a verification program that stops at a single passing specimen reproduces that exact historical pattern instead of correcting it.

A single device, however carefully measured, cannot distinguish a real, repeatable field effect from a rare favorable coincidence of geometry, film defect, and contact condition that happened to survive Steps 15–18 and produce a Step 23 pass. Only a population can make that distinction, because only a population has a rate.

## 1. Define the population and the independence it requires before building it

A population means specimens fabricated as separate instances of the full Steps 1–14 sequence, not multiple measurements of one device, and not multiple devices cut from a single sulfurization batch sharing one film-forming event. Record, for every specimen, which fabrication sub-steps were shared (a common foil lot, a common sulfurization furnace run, a common coating session) and which were independent, because a defect common to a shared sub-step can make several specimens fail or pass together for a reason that has nothing to do with the field-effect mechanism being tested.

State the population size before any specimen is measured, using a number large enough that the statistics in §4 can discriminate a real effect rate from a chance rate at a stated confidence, not a number chosen after seeing how many early specimens passed. Revising the target size upward only after an unfavorable early run is the same protocol-adaptation failure Step 21 was written to prevent, applied now at the population level instead of the single-device level.

Specimen population with shared and independent fabrication sub-steps A set of parallel fabrication lanes shows several specimens passing through the shared steps of foil sourcing, film sulfurization, and terminal coating before diverging into independent assembly and verification lanes, with shared sub-steps flagged as a correlated-failure risk distinct from the per-specimen verification arc. A POPULATION NEEDS INDEPENDENCE, NOT JUST A HEADCOUNT Shared fabrication sub-steps create correlated outcomes that a simple pass-rate count will not reveal SHARED STEPS FUNNEL INTO INDEPENDENT LANES SHARED FOIL LOT Step 3 SHARED SULFURIZATION RUN Step 8 specimen Aspecimen Bspecimen Cspecimen D independent Steps 15-23independent Steps 15-23independent Steps 15-23independent Steps 15-23 verdictverdictverdictverdict WHAT A SHARED SUB-STEP PUTS AT RISK A single contaminated foil lot or one bad sulfurization run can make several specimens fail together for a shared reason. A simple pass-rate count across such specimens is not an independent-trial count; it overstates the information in the sample. Record which lanes share which sub-step before the population is measured. The independent-count correction in §4 downweights clusters that trace to a documented shared sub-step. Specimens built from genuinely separate foil lots and separate sulfurization runs count as fully independent trials. The population size and its independence structure are both fixed in writing before the first specimen reaches Step 15.

## 2. Run every specimen through the full, unmodified Steps 15–23 arc

Each specimen receives the identical geometry inspection, film characterization, isolation check, baseline measurement, leakage measurement, frozen bias protocol, modulation test, and power-gain quantification already defined in Steps 15 through 23, with no shortcuts taken for specimens that look unpromising early and no extra scrutiny applied only to specimens that look promising early. Measuring a favorable-looking specimen more carefully than an unfavorable-looking one biases the population toward favorable verdicts before the comparison in §4 ever runs.

Carry forward every intermediate disposition, not just the Step 23 verdict. A specimen rejected at Step 18 for a conductive bridge, or at Step 20 for leakage, is a data point about the fabrication process, not a specimen to be quietly discarded and replaced without being counted. Report the full funnel: how many specimens entered Step 15, how many survived each intermediate reject point, and how many reached a Step 23 verdict at all.

## 3. Hold the frozen protocol and thresholds fixed across every specimen

The Step 21 bias list, interleaving schedule, and the $\delta I_{min}$ and $f_{leak,max}$ thresholds from Steps 19 and 20 must be the same written protocol applied to every specimen, not re-derived per specimen from that specimen's own early data. A per-specimen threshold that is quietly loosened for a device that would otherwise fail, or tightened for one that would otherwise pass by a wide margin, destroys the comparability that makes a population meaningful. If a specimen's own baseline repeatability genuinely differs enough to need a different $\delta I_{min}$, document that as a property of the specimen and carry the original population-level threshold forward as well, so both can be reported.

## 4. Report a rate with an interval, not a count of successes

Convert the per-specimen verdicts into a population-level statistic. For $n$ specimens reaching a final Step 23 disposition, with $k$ passing, report the pass rate

$$ \hat p=\frac{k}{n}, $$

together with a binomial confidence interval appropriate for small $n$ — a Wilson or Clopper–Pearson interval rather than a normal approximation, since early-stage reproducibility runs are typically small. Where §1 identified specimens sharing a sub-step, do not treat them as $n$ fully independent trials; either exclude all but one specimen per shared cluster from the rate calculation, or use a clustering-adjusted estimator, and report which choice was made and why.

State plainly what the interval does and does not support. A high $\hat p$ with a wide interval at small $n$ supports cautious optimism and a larger follow-up population, not a settled claim. A low $\hat p$, or an interval that comfortably includes rates consistent with chance alignment of the confounds Step 23 already named, supports a conclusion that the Step 23 pass on the original specimen was not representative.

Population pass rate with confidence interval and disposition funnel A funnel diagram tracks specimens from entry through each intermediate reject point to a final Step 23 verdict, beside a pass-rate plot with a binomial confidence interval compared against a chance-rate reference, leading to a disposition rule for the reproducibility claim as a whole. A RATE WITH AN INTERVAL, NOT A SINGLE SUCCESS STORY Track the full funnel from entry to verdict, then bound the pass rate honestly at this sample size DISPOSITION FUNNEL entered Step 15 survived Steps 15-18 survived Steps 19-20 reached Step 22 verdict passed Step 23 Every reject point is counted, not discarded and silently replaced. PASS RATE WITH CONFIDENCE INTERVAL pass rate, 0 to 1 specimen group chance-rate reference p-hat with interval 01 REPRODUCIBILITY DISPOSITION PASSinterval sits clearly above the chance-rate reference across independent specimensHOLDinterval is wide, straddles the reference, or independence is uncertain — enlarge the populationREJECTinterval sits at or below the chance-rate reference, or passes cluster on a shared sub-step

## 5. Define the chance-rate reference before comparing against it

The dashed reference in §4 is not an arbitrary benchmark; it must be derived from the same confounds Steps 18 through 23 were built to exclude. Estimate, from the fixture-blank and insulating-blank data already collected across Steps 18 and 20, how often a specimen with no real control-electrode effect would be expected to produce a Step 19–23 verdict pattern that looks like a pass purely from leakage variability, baseline noise, or an unflagged correlated artifact. This reference rate — not zero, and not an assumed figure — is the rate a real effect must clear.

## 6. Characterize the failing specimens, not only the passing ones

Fold the geometry, coverage, and material maps from Steps 15 through 17 back into the population analysis. If specimens that pass Step 23 share a measurable trait — a narrower minimum foil clearance, a particular film-thickness range, a specific coverage pattern — that correlation is itself a finding, and it may point toward the physical condition the field effect actually depends on, or toward a specific fabrication defect that happens to produce a pass through a path §5's reference rate does not anticipate. Report this correlation explicitly rather than treating passing specimens as an undifferentiated group.

## 7. Minimum Step 24 record

Retain:

  • the pre-declared population size, its independence structure across shared fabrication sub-steps, and the written justification for both, fixed before the first specimen reached Step 15;
  • the complete Steps 1–14 fabrication record for every specimen, with shared sub-steps flagged;
  • the complete Steps 15–23 record for every specimen, including every intermediate reject disposition, not only specimens reaching a final Step 23 verdict;
  • confirmation that the Step 19–21 thresholds and frozen protocol were applied identically across all specimens, with any documented exception stated;
  • the full disposition funnel from entry to final verdict;
  • the pass rate $\hat p$, its confidence interval and the method used, and the independence-adjustment method where shared sub-steps were present;
  • the chance-rate reference derived from Step 18 and Step 20 blank data, and the comparison against it;
  • any measured geometry, coverage, or material correlation distinguishing passing from failing specimens;
  • the final reproducibility disposition for the population as a whole.

The historical geometry and claimed effect are grounded in [J. E. Lilienfeld, US Patent 1,745,175](https://patents.google.com/patent/US1745175A/en). The requirement to establish a rate across independent trials before accepting a device-level claim, rather than generalizing from a single specimen, follows the same discipline that separates a reproducible instrumental result from an isolated favorable observation throughout the low-level measurement practice referenced in the [Keithley Low Level Measurements Handbook](https://www.tek.com/en/documents/product-article/keithley-low-level-measurements-handbook---8th-edition).

## Establish Reproducibility Across Specimens’ Place in the Process Lineage

Steps 15 through 23 built a rigorous verification arc and applied it to one specimen. Step 24 now runs that same arc, unmodified, across a pre-declared population of independently fabricated specimens, tracks every intermediate disposition rather than discarding unfavorable ones, and reports a pass rate with a confidence interval against a chance-rate reference derived from the project's own blank measurements. Only a population-level result that clears that reference, with documented independence and a characterized relationship between passing specimens and their measured geometry, coverage, and material properties, extends the Step 23 amplification verdict from one device to a claim about the fabrication process as a whole.

Take lilienfeld 1925 establish reproducibility across specimens further

Ask the copilot about this term, or have our engineers assess it against your process.