Research Design Tools: Turn a Question into a Testable Experiment - Yenra

Plan controls, experimental units, randomization, replication and blocking, with a verified fictional filter experiment and downloadable research worksheets.

Two trays of capped sample vessels sit beside a research notebook and a glass panel of planning shapes.
Conceptual illustration: organized comparisons and documented plans support an experiment; the vessels do not depict the filter-coupon example below.

A research design connects a question to observations that can answer it. Software can organize that connection, but it cannot repair a comparison in which treatment, day, operator and sample source all change together. Start by deciding what difference you want to estimate and which observations would distinguish that difference from ordinary variation.

This guide uses a small fictional materials experiment to make the decisions concrete. The same planning questions are useful across many disciplines, but ethical review, domain methods and appropriate statistical advice remain specific to the actual project.

Write an answerable question before choosing a tool

“Which filter is better?” leaves too much undefined. A more useful question is: under a specified illumination and measurement method, how much does mean light transmission differ between two filter materials? That names a comparison and an outcome. It still needs the intended population of filters, units, measurement conditions and meaningful effect size.

Record the primary outcome before collecting results. If transmission is expressed as a percentage of incident light, a difference between 40% and 45% is five percentage points. It is not a five-percent relative increase. The units of the answer are part of the design, not a formatting decision made at the end.

The NIST guide to choosing an experimental design begins with objectives and variables before selecting a design. Use that order when comparing tools: first determine the work the design must do, then choose a notebook, spreadsheet or statistical environment that can represent it.

Count independent units, not rows in a file

The experimental unit is the entity independently assigned to a treatment or condition. A measurement is an observation on that unit. Measuring the same filter coupon ten times may help characterize reading variability, but it does not create ten independently produced filter coupons.

On a small screen, scroll the table sideways to read all columns.

A planning vocabulary for the fictional filter study
ItemDefinition in this exampleWhy it matters
Experimental unitOne independently produced filter coupon.The unit represented by an independent material specimen.
Treatment or factorMaterial A versus material B.The comparison whose effect is of interest.
OutcomeLight transmission under the specified setup, in percent.The quantity that must be measured consistently.
BlockMeasurement day, with both materials measured each day.Day-to-day differences can be separated from the material comparison.
Repeated readingAnother reading of the same coupon.Useful measurement information, but not a new independent coupon.

Independence needs a physical justification. Several pieces cut from one sheet may share manufacturing variation; they should not automatically be treated as independent production replicates. If the intended conclusion concerns manufacturing batches, the design needs information from independent batches. State the level at which the experiment can support a conclusion.

Use randomization, replication and blocking for different jobs

NIST's experimental-design principles explain why randomization and replication belong in a planned experiment. Randomization reduces systematic alignment between conditions and nuisance influences. Replication supplies information about variation. Neither guarantees that an unrepresentative set of specimens represents a wider population.

Blocking deliberately groups comparable conditions. In a randomized block design, treatments are compared within blocks rather than allowing a nuisance factor to become inseparable from treatment. If every A coupon is measured on Monday and every B coupon on Tuesday, a day effect is entangled with the material difference.

For the example, allocate two independent coupons of each material to each of three days, then randomize measurement order within each day. Keep the same measurement method and document any instrument checks. This balances the comparison across days. If practical, use coded specimen labels so the person recording the reading does not know which material is expected to perform better.

Choose the number of independent units using the effect worth detecting, expected variation, desired uncertainty and available resources. Twelve coupons are used here to keep an example readable; twelve is not a universal sample-size recommendation. A pilot can help estimate variation, but should not be relabeled as definitive simply because its difference looks large.

Worked example: compare within each day

These are invented transmission measurements, not claims about real filter materials. Each cell below contains two independent coupons, each measured once.

On a small screen, scroll the table sideways to read all columns.

Fictional transmission data and within-day differences
DayMaterial A (%)Material B (%)Mean B − mean A
142, 44 → mean 4348, 50 → mean 496 percentage points
250, 52 → mean 5155, 57 → mean 565 percentage points
346, 48 → mean 4754, 56 → mean 558 percentage points

The mean of the three daily differences is (6 + 5 + 8) / 3 = 6.33 percentage points, rounded to two decimals. Because the allocation is balanced, this also equals the overall B mean of 53.33% minus the A mean of 47.00%. The day-by-day display exposes variation that the two overall means conceal.

The arithmetic does not provide a confidence interval, establish statistical significance or justify extrapolation to other wavelengths, temperatures or manufacturing lots. A formal analysis must reflect the design and its assumptions, including whether material effects vary by day. Preserve individual observations so a reviewer can evaluate those questions.

Download the fictional 12-run dataset with an illustrative within-day order (CSV) and the blank experiment-planning worksheet (CSV). The order is an example schedule, not a required order for your study. Generate and preserve the randomization appropriate to your own design before observing outcomes.

Make the plan reproducible and the deviations visible

A useful project record connects each specimen identifier to its origin, assigned condition, block, run order, measurement settings and raw file. Keep a dated version of the plan and document deviations without silently replacing the original. A failed reading and a missing specimen are different events and should have different reasons recorded.

Specify exclusion rules before examining the result wherever possible. Do not delete an inconvenient observation merely because it changes the conclusion. Investigate equipment or transcription problems, retain the original record and explain any correction. Decide how repeated readings will be summarized before allowing them to inflate the apparent sample size.

A spreadsheet can handle a small run plan; a notebook can connect analysis code with explanations; an electronic laboratory notebook can manage records and permissions. Choose based on traceability, collaboration and export needs. Software features are not proof that a design answers its question.

AI can help draft a data dictionary, check identifiers or propose analysis code. Test that code against the worked values above and deliberately altered inputs. It should report missing data and invalid units, not quietly manufacture replacements. A fluent explanation of a model is insufficient evidence that the model matches the experimental unit or dependence structure.

Related resources

Researched and updated September 6, 2026. Consult the linked primary sources for methods, evidence, and limitations.