An operational definition states exactly how a variable will be measured or manipulated in a particular study, so that anyone could observe it, score it and repeat the procedure. It turns an abstract concept into something concrete. For example, anxiety can be operationally defined as a participant's total score (0 to 21) on the GAD-7 questionnaire; stress as the change in salivary cortisol after a timed public-speaking and mental-arithmetic task; and memory as the number of words, out of 20, a participant correctly recalls after a 10-minute delay. A good operational definition names the instrument or procedure, how it is scored, and when it is measured.
Working on a psychology methods assignment or lab report? Get one-to-one support from a specialist. Get a free quote →
What is an operational definition in psychology?
Most things psychologists study cannot be seen directly. Nobody can point a ruler at anxiety, stress, intelligence or memory. These are constructs: ideas that explain behaviour and experience but have to be inferred from something observable. An operational definition bridges that gap. The APA Dictionary of Psychology (American Psychological Association, n.d.) describes it as defining a concept in terms of the operations or procedures used to measure or produce it.
It helps to separate two kinds of definition:
- Conceptual (theoretical) definition: what the construct means. Anxiety is a state of apprehension and worry about possible future threats, with physical tension.
- Operational definition: how the construct is measured or manipulated in this study. Anxiety is the participant's total score on the 7-item GAD-7, completed at the start of the session.
The conceptual definition tells readers what you are interested in; the operational definition tells them what you actually recorded. Both belong in a research report, and the operational definition should follow logically from the conceptual one.
Why operational definitions matter
- Replication. Another researcher can repeat your study only if they know exactly what you measured and how. "Participants were stressed" cannot be repeated; "participants completed the Trier Social Stress Test" can.
- Objectivity. Two observers applying the same definition should record the same result, which keeps personal judgement out of the data.
- Communication. Two studies can both claim to measure "memory" while measuring very different things. Precise definitions let readers compare findings fairly.
- Testable hypotheses. A hypothesis such as "mindfulness reduces anxiety" becomes testable only once both variables are operationalised.
- Accountability in practice. Cooper et al. (2020) point out that clinicians and teachers need explicit definitions too, even when they are not publishing: without one, they cannot apply a procedure consistently or show clients, parents and supervisors believable evidence that it is working.
How to write an operational definition in four steps
- State the construct conceptually. Write one sentence on what you mean by the construct, ideally drawing on a textbook or published theory.
- Choose the type of measure. Psychological constructs are usually measured in one of three ways: self-report (questionnaires, rating scales, interviews), behavioural (what people do, such as reaction times, errors, avoidance or words recalled) or physiological (heart rate, skin conductance, cortisol, brain activity).
- Specify the procedure and scoring. Name the instrument, the number of items or trials, the response format, how the score is calculated, the possible range, and when and where it is taken. If you use a cutoff to form groups, state it. When you observe behaviour, also say what does not count, so observers know where the boundaries are.
- Check reliability and validity. Choose a measure with published evidence that it is consistent and actually captures the construct, and report that evidence.
A useful template is: "[Construct] was operationally defined as [score, count or change] on [instrument or task], [scored how], measured [when]; [what is excluded]." For a manipulated variable: "[Construct] was manipulated by [exact procedure, duration and instructions], compared with [control condition]."
Weak: Memory was measured with a memory test.
Better: Memory was operationally defined as the number of words (0-20) correctly written down in a 2-minute free-recall test, given 10 minutes after participants studied a list of 20 unrelated nouns presented one at a time for 2 seconds each. Misspellings were accepted; words not on the list were scored as intrusions and not counted.
Operational definitions of anxiety: examples
Anxiety is one of the most frequently operationalised constructs in psychology courses, and it can be measured in all three ways. Before choosing, decide whether you mean state anxiety (how anxious someone feels right now, which can change within minutes) or trait anxiety (a stable tendency to feel anxious across situations). They call for different measures and different timing.
| Type | Operational definition | Notes |
|---|---|---|
| Self-report | Total score on the GAD-7 (7 items about the past two weeks, each scored 0-3; range 0-21) | Spitzer et al. (2006) proposed 5, 10 and 15 as cutpoints for mild, moderate and severe anxiety |
| Self-report | Anxiety subscale score on the Hospital Anxiety and Depression Scale (HADS-A; 7 items, range 0-21; Zigmond & Snaith, 1983) | Designed for medical settings; avoids physical symptoms that could reflect illness |
| Self-report (state) | Rating from 0 (not at all anxious) to 100 (extremely anxious) on a single visual analogue scale, taken immediately before and after a task | Quick and sensitive to change, but a single item has limited reliability |
| Behavioural | Number of seconds a participant stands at the front of the room before choosing to stop a speech, or the distance they approach a feared object in a behavioural approach test | Captures avoidance, a core feature of anxiety |
| Physiological | Increase in heart rate (beats per minute) or skin conductance from a 5-minute resting baseline to the anticipation period of a task | Objective, but arousal also rises with excitement or exercise |
The GAD-7 shows why the details matter. Spitzer et al. (2006) found that a score of 10 or more identified generalized anxiety disorder with good sensitivity and specificity in primary care. That makes it a screening cutoff, not a diagnosis. If your study forms a "high anxiety" group from GAD-7 scores of 10 or above, say so, and do not describe those participants as diagnosed.
Hypothesis: Students who complete a brief breathing exercise will report lower state anxiety before an oral presentation than students who read a neutral passage.
Independent variable: a 5-minute paced-breathing audio guide (6 breaths per minute) versus a 5-minute neutral reading task, randomly assigned.
Dependent variable: anxiety rating on a 0-100 scale completed immediately before the presentation.
Operational definitions of stress: examples
"Stress" can refer to the event (a stressor), the person's appraisal of it, or the body's response. A clear operational definition shows which one you mean. Stress is also unusual in that it is often used as an independent variable that researchers induce, as well as a dependent variable they measure.
| Role | Operational definition | Notes |
|---|---|---|
| Measured (self-report) | Total score on the Perceived Stress Scale (PSS; Cohen et al., 1983), which asks how unpredictable, uncontrollable and overloaded life has felt in the past month | Measures appraised stress rather than the number of events; the 10-item version is widely used |
| Measured (physiological) | Change in salivary cortisol (nmol/L) from a baseline sample to a sample taken about 20-40 minutes after a stressor begins, when responses typically peak | Timing is critical, because cortisol responds with a delay |
| Measured (physiological) | Mean systolic blood pressure or heart rate during a task compared with rest | Responds within seconds; useful alongside cortisol |
| Manipulated | Trier Social Stress Test (Kirschbaum et al., 1993): a preparation period, a 5-minute mock job interview speech and 5 minutes of mental arithmetic in front of an unresponsive panel | A standardised laboratory stressor that reliably raises cortisol |
| Manipulated | Performing serial subtraction aloud under time pressure, with errors pointed out, versus counting aloud at an easy pace | Simpler to run; still needs a matched control condition |
Dickerson and Kemeny's (2004) meta-analysis of 208 laboratory studies found that tasks combining uncontrollability with social-evaluative threat, being judged by others, produced the largest cortisol responses. That explains the design of the Trier Social Stress Test and is a good justification to cite if you use a social-evaluative stressor.
Hypothesis: Acute social stress impairs working memory.
Independent variable (stress): the Trier Social Stress Test versus a control version in which participants speak about a neutral topic and count aloud, with no audience.
Manipulation check: salivary cortisol and a 0-100 stress rating, each taken before and after the task.
Dependent variable (memory): longest backward digit span correctly repeated on two consecutive trials.
Operational definitions of memory: examples
Memory is not one thing. Psychologists distinguish short-term and working memory, long-term episodic memory, recognition and recall, and true and false memories. Your operational definition should name the kind of memory as well as the task.
| Kind of memory | Operational definition |
|---|---|
| Short-term/working memory | Longest sequence of digits a participant repeats correctly, forwards or backwards, with sequences increasing by one digit until two consecutive errors |
| Free recall | Number of words correctly recalled, in any order, out of a 20-word list after a set delay |
| Recognition | Hit rate minus false-alarm rate when old words are mixed with an equal number of new words |
| Long-term episodic memory | Number of correct answers to 10 questions about a short film, asked one week after viewing |
| False memory | Proportion of trials on which participants recall a related word that was never presented (the Deese-Roediger-McDermott, or DRM, task) |
Two published findings show how a definition shapes what a study can conclude. Cowan (2001) argues that when rehearsal and chunking are prevented, short-term storage holds only about three to five chunks, far fewer than the "seven plus or minus two" often quoted, so the way a span task controls rehearsal changes the answer it gives. In the DRM task, Roediger and McDermott (1995) presented lists such as bed, rest, awake, tired, dream and found that participants frequently recalled and "remembered" the related word sleep, which never appeared. Defined as the number of correct words recalled, memory looks accurate; defined as intrusions of related lures, the same task reveals systematic errors.
Hypothesis: Testing yourself improves long-term retention more than rereading.
Independent variable: after reading a 500-word passage once, participants spend 10 minutes either answering practice questions without feedback (retrieval practice) or rereading the passage, randomly assigned.
Dependent variable: number of correct answers (0-20) on a short-answer test about the passage taken 48 hours later.
Defining observable behaviour: lessons from applied behavior analysis
Applied behavior analysis (ABA) has the most detailed rules for operational definitions, because its data come from people observing behaviour directly. Applied Behavior Analysis (Cooper et al., 2020), the standard textbook in the field, calls the behaviour being studied the target behavior and insists it be defined clearly, objectively and concisely before any analysis begins. Baer et al. (1968) called this the technological dimension of the field: a procedure or behaviour should be described well enough that a trained reader could reproduce it from the description alone.
Function-based vs topography-based definitions
- Function-based definition: counts any response that has the same effect on the environment, whatever it looks like. Escape from a feared situation was recorded whenever the participant left the room within 10 seconds of the spider jar being uncovered. Cooper et al. (2020) recommend function-based definitions whenever possible, because they capture every form of the behaviour, are usually simpler, and are easier to record reliably.
- Topography-based definition: counts responses by their shape or form. Nail biting was recorded each time a fingernail made contact with the teeth. Use this when you cannot observe the outcome directly, or when the outcome can be produced by other events or by unwanted variations of the behaviour.
Cooper et al. (2020) also warn against judging behaviour by form alone. The same movements, such as rubbing an eye or squeezing someone's arm, can be a problem in one context and harmless or useful in another, so the definition should be anchored in what the behaviour does, not only in what it looks like.
Choosing what to measure: count, rate, duration and latency
Behaviour happens over time, so you can measure it by how often it occurs, how long it lasts and when it happens. Your operational definition should state which of these dimensions you are recording:
| Dimension | What it records | Example |
|---|---|---|
| Count | Number of times the behaviour occurs | Number of reassurance-seeking questions asked during a 20-minute session |
| Rate | Count per unit of time, so sessions of different lengths can be compared | Nail-biting episodes per 10 minutes of study time |
| Duration | Time from the start to the end of each occurrence, or in total | Total seconds spent looking away from the audience during a 3-minute speech |
| Latency | Time from a cue to the start of the response | Seconds between the instruction "begin your speech" and the first word spoken |
| Inter-response time | Time between the end of one response and the start of the next | Seconds between successive words recalled in a free-recall test |
Operationalising independent and dependent variables
The table below brings the three constructs together and shows a full operational definition for both variables in each study.
| Research question | Independent variable (operationalised) | Dependent variable (operationalised) |
|---|---|---|
| Does exercise reduce anxiety? | 20 minutes of cycling at moderate intensity versus 20 minutes of seated rest | Change in 0-100 state anxiety rating from before to after the session |
| Does sleep loss raise perceived stress? | One night restricted to 4 hours in bed versus 8 hours, in a lab | PSS total score the following evening (wording adapted to "today") |
| Does background music impair recall? | Studying a word list with lyrical music, instrumental music or silence | Number of words recalled (0-30) after a 5-minute distractor task |
| Are first-year students more anxious than final-year students? | Year of study (quasi-independent; not assigned) | GAD-7 total score in week 4 of semester |
The last row uses a characteristic the researcher cannot assign, which makes it a quasi-independent variable and the study non-experimental or quasi-experimental. Our guide to quasi-experimental design explains what that means for causal conclusions. How you operationalise the dependent variable also decides the analysis: a 0-21 total score is usually analysed as continuous data with a t-test or ANOVA, while a cutoff that sorts people into "anxious" and "not anxious" produces categories that call for a chi-square test. See which statistical test to use for how that choice changes the test.
The independent variable needs the same care. Gresham et al. (1993) argued that a treatment or manipulation should be described as explicitly as the outcome, and suggested specifying it on four dimensions: verbal (what is said), physical (what is done), spatial (where things and people are) and temporal (when and for how long). A precise definition of the procedure also lets you check treatment integrity, or procedural fidelity: whether each session was delivered as written.
Temporal: immediately after the participant finishes reading the word list,
Verbal: the researcher says, "Now count backwards in threes from 300, out loud, until I say stop,"
Physical: starts a 30-second timer and corrects any error by repeating the last correct number,
Spatial: while seated across the table from the participant, with the word list turned face down.
Good vs weak operational definitions
Hawkins and Dobes (1977, as cited in Cooper et al., 2020) set out three characteristics of a good behavioural definition that are still treated as the standard. A definition should be:
- Objective: it refers only to what can be observed, and translates inferential phrases such as "showing interest" or "feeling hostile" into observable terms.
- Clear: it is readable and unambiguous, so an experienced observer could paraphrase it accurately.
- Complete: it sets the boundaries of what counts and what does not, so observers are left with little to judge for themselves.
A quick three-question test from Morris (1985, as cited in Cooper et al., 2020) asks: Can you count how many times the behaviour occurs in a set period, or time how long it takes? Would a stranger know exactly what to look for? Can the behaviour be broken into smaller, more observable parts? A good definition gets "yes", "yes" and "no". The pairs below show common weaknesses and how to fix them.
| Weak | Problem | Improved |
|---|---|---|
| Anxiety was measured by how nervous participants seemed. | Subjective; no instrument or scoring | Two trained observers counted visible anxiety behaviours (fidgeting, gaze aversion, speech pauses over 3 seconds) during a 3-minute speech, using a written coding scheme |
| Stress was high exam stress. | Names a situation, not a measurement | PSS-10 total score completed during the week before final exams |
| Memory was tested with a word list. | No list length, timing or scoring | Number of words recalled out of 15 after a 30-second distractor task of counting backwards in threes |
| Participants were given a stressful task. | Not replicable; no control condition | 5 minutes of serial subtraction aloud from 1,022 in steps of 13 before an evaluator, restarting after each error, versus counting aloud from 1 |
| Happy participants were compared with sad participants. | No inducing procedure or check | Mood was induced with a 4-minute film clip (comedy versus sad scene); a 1-7 mood rating after the clip served as the manipulation check |
Reliability, validity and the limits of any single definition
An operational definition can be perfectly precise and still miss the construct. Counting how often someone checks their phone during an exam is easy to measure, but it may reflect boredom rather than anxiety. Two questions guard against this:
- Reliability: does the measure give consistent results? Look for internal consistency (for example, Cronbach's alpha) for questionnaires, test-retest reliability for stable traits, and inter-rater agreement when people code behaviour.
- Inter-observer agreement (IOA): when behaviour is observed, have two people record the same sessions independently. The simplest index, total count IOA, divides the smaller count by the larger and multiplies by 100: if one observer records 18 episodes and the other 20, agreement is 90%. Cooper et al. (2020) note that a mean of at least 80% is the usual convention in applied behavior analysis, while stressing that the right level depends on what is being measured. Low agreement is often a sign that the definition itself is unclear.
- Construct validity: does the measure actually capture the construct? Cronbach and Meehl (1955) argued that this is established by showing that scores relate to other variables in the way theory predicts: higher with other anxiety measures, for example, and lower with measures of calm.
For observed behaviour, Cooper et al. (2020) add a practical test of validity: the definition should capture every aspect of the behaviour that the person raising the concern cares about, and nothing else. A definition of classroom disruption that counts only shouting, when the teacher is concerned about students leaving their seats, is reliable but not valid.
Flake and Fried (2020) describe questionable measurement practices that weaken research: not reporting which measure was used, changing items or scoring without saying so, and creating a new scale on the spot without evidence that it works. The safest route for a course project is to use an established instrument, report its source, and keep any changes to a minimum, explaining each one.
No single operation captures a construct completely, a problem sometimes called mono-operation bias. Measuring stress with both cortisol and the PSS, or anxiety with both a questionnaire and heart rate, makes conclusions more convincing when the measures agree, and informative when they do not.
Reporting operational definitions in an APA-style Method section
In a lab report or journal article, operational definitions appear mainly in the Method section, under Materials (or Measures) and Procedure. The APA Journal Article Reporting Standards (JARS; Appelbaum et al., 2018) ask authors to define every primary and secondary measure, describe how it was collected, and report evidence of its reliability and validity, including reliability estimates calculated in their own sample where possible.
Measures. Anxiety. State anxiety was measured with the Generalized Anxiety Disorder 7-item scale (GAD-7; Spitzer et al., 2006). Participants rated how often they had been bothered by seven problems over the past two weeks on a scale from 0 (not at all) to 3 (nearly every day). Item scores were summed (range 0-21), with higher scores indicating greater anxiety. Internal consistency in the present sample was good (α = .88).
Recall. Memory was operationally defined as the number of words (0-20) correctly recalled in a 2-minute written free-recall test administered 10 minutes after the study list.
Notice that the paragraph names the instrument and its original source, gives the response scale and scoring, states the direction of scores and reports reliability. The alpha value here is illustrative; report the value from your own data. When you move on to the Results, our guide to reporting statistics in APA 7 shows how to present the analyses, and if you are planning a study from scratch, how to write a research proposal walks through defining variables at the design stage.
If you are working through a research methods course and want someone to check your variables, hypotheses or analysis plan, our psychology research methods and statistics support pairs you with a specialist who explains the reasoning so you can apply it yourself.
Frequently asked questions
What is an operational definition in psychology?
It is a precise statement of how a variable is measured or manipulated in a particular study, such as the instrument, task, scoring and timing, so that others can observe the same thing and repeat the study.
What is an example of an operational definition of anxiety?
Anxiety can be operationally defined as a participant's total score (0-21) on the GAD-7 questionnaire, as a 0-100 rating of how anxious they feel right now, or as the increase in heart rate from rest to the moment before giving a speech.
How do you operationally define stress?
As a measured variable, stress can be the total score on the Perceived Stress Scale or the change in salivary cortisol after a task. As a manipulated variable, it can be induced with a standardised procedure such as the Trier Social Stress Test, compared with a non-stressful control task.
How do you operationally define memory?
Name the kind of memory and the task, for example the number of words correctly recalled out of 20 after a 10-minute delay, the longest digit span repeated correctly, or the hit rate minus false-alarm rate on a recognition test.
What is the difference between a conceptual and an operational definition?
A conceptual definition explains what a construct means in theory, such as anxiety being apprehension about future threats. An operational definition explains how it is measured in a specific study, such as a GAD-7 score.
Does the independent variable need an operational definition?
Yes. The independent variable needs a precise description of each condition, including the procedure, duration and instructions, and the dependent variable needs a precise description of how it is measured and scored.
What makes a good operational definition?
It is objective, clear and complete: it refers only to observable or recorded events, others can read and apply it without confusion, and it states what does and does not count. It should also name the instrument or procedure, how it is scored, and use a measure with evidence of reliability and validity.
What is the difference between a function-based and a topography-based definition?
A function-based definition counts any response that has the same effect on the environment, such as leaving the room within 10 seconds of seeing a feared object. A topography-based definition counts responses by their form, such as each time a fingernail touches the teeth. Behavior analysts prefer function-based definitions when the outcome can be observed.
Sources
- American Psychological Association. Operational definition. APA Dictionary of Psychology
- Cooper JO, Heron TE, Heward WL. Applied Behavior Analysis. 3rd ed. Hoboken, NJ: Pearson; 2020 (chapters 3-5 and 10)
- Baer DM, Wolf MM, Risley TR. Some current dimensions of applied behavior analysis. J Appl Behav Anal 1968;1:91-97
- Gresham FM, Gansle KA, Noell GH. Treatment integrity in applied behavior analysis with children. J Appl Behav Anal 1993;26:257-263
- Cronbach LJ, Meehl PE. Construct validity in psychological tests. Psychol Bull 1955;52:281-302
- Flake JK, Fried EI. Measurement schmeasurement: questionable measurement practices and how to avoid them. Adv Methods Pract Psychol Sci 2020;3:456-465
- Spitzer RL, Kroenke K, Williams JBW, Löwe B. A brief measure for assessing generalized anxiety disorder: the GAD-7. Arch Intern Med 2006;166:1092-1097
- Zigmond AS, Snaith RP. The Hospital Anxiety and Depression Scale. Acta Psychiatr Scand 1983;67:361-370
- Cohen S, Kamarck T, Mermelstein R. A global measure of perceived stress. J Health Soc Behav 1983;24:385-396
- Kirschbaum C, Pirke KM, Hellhammer DH. The 'Trier Social Stress Test': a tool for investigating psychobiological stress responses in a laboratory setting. Neuropsychobiology 1993;28:76-81
- Dickerson SS, Kemeny ME. Acute stressors and cortisol responses: a theoretical integration and synthesis of laboratory research. Psychol Bull 2004;130:355-391
- Roediger HL, McDermott KB. Creating false memories: remembering words not presented in lists. J Exp Psychol Learn Mem Cogn 1995;21:803-814
- Cowan N. The magical number 4 in short-term memory: a reconsideration of mental storage capacity. Behav Brain Sci 2001;24:87-114
- Appelbaum M, Cooper H, Kline RB, et al. Journal article reporting standards for quantitative research in psychology: the APA Publications and Communications Board task force report. Am Psychol 2018;73:3-25
