Speech-language pathologists collect data constantly.
A clinician may record 18 correct productions out of 25 trials for initial /r/, note that a child answered four of six inferential questions correctly, document the number of independent AAC requests during a classroom routine, or rate the level of cueing needed to produce a narrative.
Those numbers are useful.
But data collection becomes clinically valuable only when the clinician can answer:
- What exactly was measured?
- Under what conditions?
- How much support was provided?
- Is performance changing over time?
- Does the change generalize beyond the treatment task?
- What should happen next because of these data?
That is the difference between recording therapy performance and using data for clinical decision-making.
The American Speech-Language-Hearing Association's evidence-based practice framework integrates clinical expertise, external research evidence, internal evidence generated through clinical data and observation, and client, patient, or caregiver perspectives.
Internal clinical evidence matters because treatment research tells clinicians what has worked across groups of participants. Session and outcome data help answer a different question:
Is this intervention producing useful change for this individual?
A strong speech-therapy data system should therefore connect:
goal → treatment condition → session data → progress trend → functional outcome → clinical decision
Technology can make that chain easier to manage. It cannot determine what should be measured or what the data mean.
What Is Speech Therapy Data Collection?
Speech therapy data collection is the systematic recording of information about a client's communication, cognition, speech, language, voice, fluency, swallowing, or related functional performance.
The data may be quantitative.
For example:
- 16/20 accurate productions;
- 80% independent responses;
- 7 spontaneous communication initiations;
- 102 words correct per minute; or
- 4 communication repairs across 5 breakdown opportunities.
Or qualitative.
For example:
- The student generated relevant inferences but required a visual organizer to cite supporting evidence.
- Speech intelligibility decreased substantially during unstructured peer conversation despite high accuracy during sentence-level drill.
- The client independently used AAC during structured therapy but required partner prompting during community interaction.
Both kinds of information can contribute to clinical reasoning.
A 2022 scoping review of clinical evidence in speech-language pathology identified practice-based evidence—the local clinical data generated during care—as a distinct source of evidence that can contribute to decision-making alongside research and clinician expertise.
Data Collection, Assessment, Progress Monitoring, and Documentation Are Different
These terms are often used as if they refer to the same process. They answer different questions.
| Process | Primary question |
|---|---|
| Assessment | What does the client's communication profile look like? |
| Baseline measurement | Where is the client starting on this specific target? |
| Session data collection | How did the client perform during today's treatment conditions? |
| Progress monitoring | Is performance changing over time? |
| Outcome measurement | Has intervention produced a clinically or functionally important change? |
| Documentation | What service was provided, why was it skilled, and what happened? |
A standardized language assessment administered every session would not constitute good progress monitoring.
Likewise, collecting 50 articulation trials during one session does not tell us whether the student has improved over three months.
The measure has to match the clinical question.
Assessment Provides the Clinical Starting Point
Assessment is broader than treatment data collection.
Depending on the referral question, an evaluation may include standardized measures, criterion-referenced measures, language sampling, speech-sound analysis, dynamic assessment, caregiver or teacher report, observation, oral-mechanism examination, hearing-related information, AAC assessment, or functional communication measures.
These data help determine the nature of the communication difficulty, strengths, barriers, functional effects, and possible intervention priorities.
A treatment data sheet should not substitute for a comprehensive evaluation when one is required.
Likewise, standardized assessment should not substitute for direct measurement of the treatment target.
Baseline Data Tell You Where the Goal Begins
Before a clinician can interpret progress, they need a defensible starting point.
Consider this goal:
The student will produce initial /r/ at the word level with 80% accuracy.
A baseline of 72% with one verbal cue creates a very different treatment problem from 5% despite maximal modeling.
The target may look identical. The starting conditions are not.
A useful baseline should therefore document the target, task level, stimulus characteristics, accuracy, cueing, and relevant context.
For example:
Initial /r/ in familiar single words: 6/20 independently; 11/20 with one verbal/visual cue.
That gives later data meaning.
Session Data Answer a Narrow Question
Session data tell the clinician how a client performed under the conditions of that session.
Suppose a student produces /s/ with 90% accuracy in imitated sentences and 45% accuracy in spontaneous conversation.
Those are not contradictory data. They describe two different levels of task complexity.
Likewise, 80% accuracy with a written visual support is not equivalent to 80% independently.
A good data system preserves those distinctions. Otherwise, a graph may appear to show progress when the treatment conditions actually changed.
Always Record the Level of Support When It Affects Performance
Accuracy alone can be misleading.
Consider three sessions:
| Session | Accuracy | Support |
|---|---|---|
| 1 | 60% | Independent |
| 2 | 80% | Moderate verbal cues |
| 3 | 90% | Direct imitation |
The numbers appear to improve. Independence may actually have decreased.
If prompting is clinically relevant, data should capture it.
Possible categories include independent, indirect cue, one verbal cue, visual cue, model, moderate prompting, and maximum support.
The categories should be defined consistently.
Define What Counts as Correct Before Collecting Data
Two clinicians should ideally score the same target in approximately the same way. That requires an operational definition.
For articulation, does a distorted but recognizable sound count as correct?
For grammar, does the student have to produce the target morpheme spontaneously?
For inferencing, must the answer be plausible? Must the student cite evidence?
For pragmatic language, what qualifies as a successful repair?
For AAC, does a prompted selection count the same as a spontaneous message?
Without defined scoring rules, percentages can look precise while representing inconsistent judgments.
Percentage Accuracy Is Useful—but It Is Not the Only Data Type
Speech-language pathology often defaults to 80% accuracy for almost every goal.
Some targets fit percentage data well. Others do not.
Percentage or Correct/Incorrect Trials
Useful for speech sounds, grammatical structures, word identification, following specific directions, and phonological-awareness items.
Frequency
Useful for communication initiations, stuttering events under defined conditions, independent AAC messages, requests for clarification, and communication repairs.
Duration
Useful for sustained phonation, periods of engagement under certain functional goals, and time spent using a strategy.
Latency
Useful for response initiation, word retrieval, and task initiation when clinically relevant.
Rating Scales or Rubrics
Useful for narratives, conversation, writing, voice quality, functional communication, and discourse organization.
Level of Cueing
Useful across many treatment domains.
The measurement method should fit the construct.
“80% Accuracy” Does Not Automatically Mean Mastery
A student may produce a target correctly in 8 of 10 trials.
But were the same ten words practiced repeatedly? Were responses imitated? Were cues provided? Was the environment quiet? Did the skill occur in conversation? Was performance stable over time?
A stronger interpretation might be:
80% accuracy in novel words across three probes with no more than one indirect cue.
Now the number has context.
Progress Monitoring Is About the Trend
Progress monitoring asks whether performance changes over repeated observations.
One strong session does not establish progress. One poor session does not establish regression.
The clinician needs enough comparable data to see a trajectory.
| Week | Initial /r/ word-level accuracy |
|---|---|
| 1 | 25% |
| 2 | 35% |
| 3 | 40% |
| 4 | 55% |
| 5 | 65% |
| 6 | 75% |
The pattern is more informative than any one point.
The relevant questions become:
- Is the slope adequate?
- Is performance stable?
- Does accuracy improve only with cueing?
- Is the student approaching the goal?
- Has progress plateaued?
- Should task complexity increase?
School-Based Progress Monitoring Has a Specific IEP Function
For students receiving services under IDEA, the IEP must specify how progress toward annual goals will be measured and when progress reports will be provided to parents.
This means the data system should allow the SLP to answer:
How is the student progressing toward the actual annual goal?
If the IEP goal addresses conversational /r/, collecting only isolated-word drill data is insufficient.
If the goal addresses communication repair, documenting generic “social skills participation” does not directly measure it.
The data should map back to the goal behavior.
Goal and Data Alignment Is One of the Most Important Principles
Consider this goal:
When a listener indicates misunderstanding, the student will independently repair the message in 4 of 5 opportunities.
Appropriate data would identify five breakdown opportunities, three independent repairs, and two repairs after a prompt.
Student participated well in social-language activity may belong in a note. It cannot establish progress toward the goal.
Probe Data and Treatment Data Serve Different Purposes
Clinicians sometimes score every practice trial and treat the result as progress data. That can overestimate learning.
If a student receives model → cue → feedback → repetition throughout therapy, performance during that practice is partly a measure of treatment support.
A probe can provide cleaner information.
Treatment: 30 /r/ words with cueing and feedback.
Probe: 10 novel /r/ words without direct models.
The probe better answers: Can the student produce the skill under less-supported conditions?
Both data types are useful. They should not necessarily be interpreted the same way.
Generalization Should Be Measured Separately
A student may master a skill during structured therapy and still not use it outside the treatment context.
For articulation:
word → phrase → sentence → reading → structured conversation → spontaneous conversation
For language:
picture description → structured response → curriculum passage → classroom discussion
For AAC:
therapy room → classroom → peer interaction → home or community
For pragmatic language:
role-play → clinician conversation → peer interaction → natural environment
Generalization data tell us whether the intervention is changing real communication rather than only treatment-task performance.
Outcome Measurement Goes Beyond Goal Accuracy
Goal attainment matters. Functional change matters too.
ASHA's National Outcomes Measurement System uses Functional Communication Measures to describe changes in communication or swallowing ability over time.
ASHA notes that standardized assessments may show improvement in specific skills but do not always capture how those improvements translate into school, work, social interaction, or other everyday activities.
This distinction is crucial.
A client may improve from 50% to 90% on a structured language task.
But does the client participate more successfully in class, communicate more independently, require fewer partner prompts, understand conversations better, return to work, eat safely, or advocate for themselves?
Functional outcome measures help answer these questions.
Patient-, Student-, and Caregiver-Reported Outcomes Add Another Layer
The clinician is not the only person who can describe change.
Client and caregiver perspectives are part of evidence-based practice.
Patient-reported outcome measures can capture experiences such as communication participation, cognitive difficulty, swallowing impact, confidence, fatigue, or perceived communication effectiveness.
ASHA's NOMS system allows optional patient-reported outcomes alongside clinician-administered functional measures.
Research in speech-language pathology has argued that patient-reported outcomes remain underused despite their importance to person-centered and evidence-based care.
A client's perception does not replace objective measurement. It answers a different question.
Formal and Informal Measures Should Be Used for Different Purposes
Formal standardized measures provide useful information when the test is valid for the individual, standard administration is preserved, normative comparison is relevant, and the construct matches the clinical question.
Informal or criterion-referenced measures may be more useful for specific therapy targets, repeated probes, functional communication, generalization, curriculum-linked performance, or highly individualized goals.
The strongest data system often includes both.
For example:
- standardized assessment → broader language profile;
- criterion-referenced probe → current grammar target;
- language sample → spontaneous use; and
- teacher report → classroom impact.
Together they provide a more complete clinical picture.
Dynamic Assessment Produces a Different Kind of Data
Dynamic assessment asks how performance changes when support or teaching is introduced.
Rather than asking only Can the child do this?, the clinician asks How does the child respond to teaching?
Relevant observations may include amount of cueing required, rate of learning, ability to transfer a strategy, responsiveness to modeling, and independence after support is withdrawn.
These data are particularly valuable when static standardized scores do not fully explain learning potential or when cultural and linguistic factors complicate norm-referenced interpretation.
Data Collection Should Match the Treatment Approach
Different interventions require different data.
Motor-Based Speech Treatment
Track accuracy, cueing, phonetic context, practice level, and generalization.
Phonological Treatment
Track contrast accuracy, pattern change, untreated generalization, and phonemic inventory where relevant.
Language Treatment
Track target form, accuracy, independence, linguistic context, and generalization to connected language.
Fluency
Depending on the treatment model, measures may include frequency, severity, physical tension, avoidance, communication attitudes, participation, or strategy use.
A single percent-fluent score is rarely sufficient for the entire clinical picture.
AAC
Track more than device-button accuracy.
Useful measures may include communicative functions, independence, communication partners, settings, repair, message complexity, and access efficiency.
Pragmatic or Social Communication
Measure actual communication behaviors, partners, context, independence, repair, self-advocacy, topic management, or functional participation.
The data architecture should follow the clinical construct.
Collect Enough Data to Make a Decision—Not Data for Its Own Sake
More data are not automatically better.
Recording every production in a 45-minute session may be useful during highly structured motor practice. It may be unnecessary during a narrative-language session.
A clinician should ask:
What decision will these data help me make?
If there is no answer, the measure may be unnecessary.
Possible decisions include increasing task difficulty, fading cues, continuing the target, changing the treatment procedure, probing generalization, revising the goal, discharging the goal, increasing support, or referring for additional assessment.
Data should support action.
Sampling Can Reduce Data Burden
Clinicians do not always need to score every opportunity.
A structured sample may be sufficient.
For example, score the first 10 opportunities, collect a five-minute conversational sample, use a defined probe at the end of treatment, or rotate detailed data across goals.
The sample needs to be representative enough for the clinical question.
A small systematic probe may provide better information than a large set of inconsistently scored observations.
Data Reliability Matters
Speech-language data are sometimes treated as inherently objective because a percentage is generated.
But scoring can be subjective.
Did the student produce a correct /r/? Was the narrative response relevant enough? Did the client independently initiate, or did the clinician's pause function as a cue? Did a stuttering event occur?
Inter-rater reliability—the extent to which two trained observers score the same behavior consistently—can be important in assessment, research, supervision, and some clinical contexts.
Even when formal reliability calculations are not practical, clinicians can improve consistency by defining targets, using scoring examples, calibrating with colleagues, recording samples when permitted, and keeping cue definitions stable.
Speech Therapy Data Sheets Should Reduce Ambiguity
A good data sheet should make it easy to see the goal, target, number of opportunities, accuracy, cueing, treatment context, and notes relevant to interpretation.
For example:
| Target | I | VC | Model | Incorrect |
|---|---|---|---|---|
| Initial /r/ | 7 | 5 | 3 | 5 |
| Final /r/ | 3 | 4 | 6 | 7 |
Where I means independent, VC means one verbal or visual cue, and Model means direct model.
This provides more information than /r/ = 73% because it distinguishes independence from supported performance.
Rubrics Are Often Better Than Percentages for Complex Language
Narrative, discourse, writing, conversation, and pragmatic language frequently contain multiple dimensions.
A narrative rubric might score character, setting, problem, event sequence, causal links, resolution, and cohesion.
A social-communication rubric might score response relevance, repair, listener information, topic management, and self-advocacy.
A voice measure might combine perceptual and instrumental information.
The rubric should define the scoring levels clearly enough that change can be tracked over time.
Session Notes and Data Sheets Are Not the Same Document
A data sheet may contain 12/15 correct productions, three spontaneous repairs, one verbal cue, and two indirect cues.
The session note communicates the clinical meaning of those data.
For example:
The student produced initial /r/ in novel single words with 80% accuracy independently, increasing from 55% on the previous comparable probe. Accuracy decreased to 50% in self-generated phrases. Treatment focused on transferring the established production from single words to phrase-level contexts.
That note tells another clinician what was treated, how the student performed, what changed, and why the next treatment step makes sense.


.png)
.png)