Guides/Interobserver agreement (IOA) in ABA

Interobserver agreement (IOA) in ABA

Interobserver agreement quantifies how closely two independent observers agree when scoring the same behavior at the same time. It is the field's core data-quality check, and it is the one number that determines how much weight every other number deserves: a treatment decision is only as trustworthy as the reliability of the data underneath it.

Reviewed August 2026

What IOA does and does not tell you

IOA measures agreement, not accuracy. Two observers can agree almost perfectly and both be wrong, usually because the operational definition they share is loose enough to be applied consistently but incorrectly. High agreement means the data are reproducible; it does not certify that the right thing was measured.

That is why low agreement is diagnostic rather than merely disappointing. Poor IOA almost always points at the definition, the training, or the measurement system — an operational definition with a judgement call in it, an observer who was never calibrated, or a partial-interval procedure applied to a behavior with ambiguous boundaries. Fixing the definition is the remedy; collecting more data under the same definition is not.

The two methods used most in practice

There are many published agreement calculations. Two cover the overwhelming majority of routine clinical use, and they differ in what they treat as a unit of comparison.

Total count IOA

smaller count ÷ larger count × 100

Use when: Each observer reports a single tally for the session, such as a count of a behavior.

Two counts of zero are treated as perfect agreement — both observers saw nothing, and that is a genuine agreement rather than an undefined result.

Exact agreement IOA — also called trial-by-trial, or point-by-point

matching units ÷ total units × 100

Use when: Each observer scores an aligned sequence, such as per-trial outcomes or per-interval marks.

Positions scored differently — or scored by only one observer — count as disagreements, so an incomplete second record lowers agreement rather than being quietly ignored.

Where the 80% convention comes from

An agreement level of at least 80% is the conventional minimum for acceptable reliability in applied behavior analysis, and it is widely used as the working threshold in practice, in supervision, and in what payers and accreditation reviewers expect to see measured. It is a convention rather than a statutory rule: it is not defined in regulation, and it should not be treated as a certification a dataset either passes or fails.

Treat it as a floor with context, not a target. Agreement of 82% on a well-defined discrete behavior may be weak for that behavior, while the same number on a difficult continuous behavior may be genuinely good. What matters more than any single check is the pattern: whether agreement is stable across observers, whether the same target keeps falling short, and whether reliability was measured often enough for the answer to mean anything.

How often to run reliability checks

Reliability has to be sampled across the conditions the data are collected in, not concentrated where it is convenient. Checks run only by one supervisor, only in clinic, or only on cooperative days will overstate agreement for the whole dataset. Spreading checks across observers, settings, and times of day is what makes the resulting figure representative.

A single check is not a reliability program. Because agreement drifts as staff turn over and definitions get informally reinterpreted, reliability is measured periodically for the life of a target — which in turn means it needs to be cheap enough to actually do, and recorded against the session and target it belongs to rather than in a separate spreadsheet.

How Cadence computes IOA

Cadence computes both methods in-product and stores each check with the session and target it belongs to, along with the raw inputs — so any result can be recomputed deterministically rather than taken on faith. Results are evaluated against the conventional 80% threshold and reported to one decimal place.

Because the checks live on the clinical record, reliability rolls up: mean agreement across checks, the proportion meeting the threshold, the lowest single agreement, and how many checks fell short. That summary is what makes a reliability program auditable at the caseload level instead of one check at a time.

FAQ

Common questions

How do you calculate interobserver agreement (IOA)?
It depends on the data. For count data where each observer reports one tally per session, total count IOA divides the smaller count by the larger and multiplies by 100. For aligned sequences such as per-trial or per-interval scores, exact agreement IOA divides the number of positions both observers scored identically by the total number of positions and multiplies by 100. Exact agreement is the more conservative of the two.
What is an acceptable IOA percentage in ABA?
At least 80% agreement is the conventional minimum for acceptable reliability, and it is the threshold most commonly used in clinical practice and supervision. It is a professional convention rather than a regulatory requirement, so it should be read as a floor to interpret in context — the difficulty of the behavior and the consistency of agreement across observers matter as much as any single figure.
What is the difference between total count and exact agreement IOA?
Total count IOA compares two session totals and asks how close they are, so two observers can reach high agreement while disagreeing about which specific instances occurred. Exact agreement, also called trial-by-trial or point-by-point IOA, compares observers position by position, so a disagreement about any single trial lowers the score. Exact agreement is stricter and is the appropriate choice whenever the data are scored as an aligned sequence.
Does IOA measure whether the data are accurate?
No. IOA measures whether two observers agree, which establishes that the data are reproducible, not that they are correct. Two observers sharing a flawed operational definition can agree closely and both be measuring the wrong thing. Persistently low agreement usually indicates a problem with the operational definition, observer training, or the measurement system rather than with the observers.
How often should IOA be measured?
Periodically for the life of a target, sampled across the observers, settings, and times of day that the data are actually collected in. Concentrating checks where they are convenient — one supervisor, one setting, cooperative sessions — overstates agreement for the dataset as a whole. Agreement also drifts as staff change and definitions get informally reinterpreted, so a single early check does not stand in for an ongoing reliability program.
Keep reading

More reference guides

The eight ABA measurement systemsWhat each of the eight ABA measurement systems measures, when to use it, and the metric it produces.

Data collection that holds up

Cadence captures every ABA measurement system natively, measures interobserver agreement in-product, and drafts the billable note from the data you just committed — for a licensed clinician to review and sign.

Start free trialSee the product