Generative AI can produce a list of /s/ words in seconds.
It can write a pirate story containing dozens of /r/ targets, create minimal-pair cards, generate sentences containing consonant clusters, turn a child's favorite topic into a speech-sound game, and rewrite an activity for different ages or levels of linguistic complexity.
That makes tools such as ChatGPT useful for one particular part of speech-language pathology:
creating and adapting treatment materials after the clinician has already made the clinical decisions.
The key phrase is "after the clinician has made the clinical decisions"
ChatGPT does not diagnose a speech sound disorder. It cannot independently determine whether a child's error is articulatory, phonological, motor based, inconsistent, developmental, or influenced by another language or dialect. It cannot decide whether /r/, /k/, a consonant cluster, or a phonological contrast should be treated first without accurate assessment information and clinical interpretation.
And a polished AI-generated activity is not automatically an evidence-based intervention.
ASHA's current guidance on generative AI states that AI technologies may support efficiency and innovation in clinical practice, but clinicians remain ethically responsible for evaluating the technology they use. ASHA specifically cautions that AI cannot replace the SLP and that claims about an AI tool's clinical efficacy need supporting evidence.
That is particularly relevant in speech sound disorders, where target selection can substantially change what the child learns.
ChatGPT is best used as:
a material-generation assistant inside an evidence-based treatment plan — not as the treatment plan itself.
What Are Speech Sound Disorders?
Speech sound disorders, or SSDs, include difficulties involving the perception, motor production, or phonological representation of speech sounds.
A child might demonstrate a relatively isolated distortion, such as a lateral /s/.
Another may systematically replace sounds produced in the back of the mouth with sounds produced in the front, as in: tea for key
Another may collapse several different adult phonemes onto one production.
Another may have inconsistent whole-word productions.
ASHA notes that the traditional distinction between articulation disorders and phonological disorders remains clinically useful, but individual children can demonstrate both types of errors. Those errors may require different treatment approaches.
Before materials are generated, the SLP needs to know what the error represents.
Assessment Comes First
Speech-sound treatment begins with clinical assessment, not content generation.
Depending on the child, an SLP may examine: single-word productions; connected speech; phonetic inventory; phonological patterns; consistency; stimulability; intelligibility; oral structures and function; speech perception; phonological awareness; hearing; and the child's communication across languages and dialects.
ASHA's current SSD guidance emphasizes that treatment selection depends on factors such as age, error type, severity, intelligibility, and the child's broader speech system.
Only after those variables are defined does an AI-generated word list become clinically useful.
For example, asking ChatGPT for 30 words containing /k/ initial position may be appropriate if the clinician has decided that /k/ production is the actual target.
It is much less useful if the child has an extensive velar-fronting pattern that would be better addressed through a phonological contrast approach.
The words may be correct. The treatment logic may not be.
The “Five Main Approaches” and Use of AI
The five categories in this framework come from Alan Kamhi's 2006 discussion of treatment decisions for children with speech sound disorders.
Kamhi described five theoretical perspectives:
But Kamhi explicitly framed these as perspectives affecting goal selection, not five separate treatments.
That means:
minimal pairs are not simply “the bottom-up approach.”
storybook practice is not inherently “the language-based approach.”
mixing several fun activities does not create a broad-based treatment model.
And:
choosing /r/ simply because it develops later does not constitute the complexity approach.
Current SSD practice contains a lot more specific interventions, including minimal oppositions, maximal oppositions, multiple oppositions, Cycles, core vocabulary, integrated phonological awareness, motor-based treatment, Speech Motor Chaining, and others.
The five-perspective framework is still useful because it helps clinicians think about why a particular target was selected.
ChatGPT can then help create materials that fit that decision.
1. The Normative Perspective
A normative perspective considers the typical developmental acquisition of speech sounds when selecting treatment targets.
A clinician using this perspective may prioritize sounds that would ordinarily be expected earlier in development before sounds that typically stabilize later.
Developmental norms can be clinically useful. They can help answer questions such as:
Is this error still commonly observed at this age?
Which absent sounds might ordinarily emerge earlier?
Is the child's phonetic inventory broadly age-appropriate?
It's important to consider that developmental order is one target-selection strategy, not a universal rule.
ASHA currently lists developmental target selection alongside several alternatives, including functional importance, effect on intelligibility, complexity, or systemic approaches.
A clinician may therefore deliberately choose a later-developing sound if another rationale supports broader change.
How ChatGPT Can Help With a Normative Target
Once the SLP has selected the sound, AI is well suited to generating developmentally appropriate practice material.
Instead of asking: “What sounds should I treat in this 4-year-old?”
the clinician retains responsibility for the target and AI generated material.
Better Prompt: Preschool Word Practice
“Create 20 familiar words for a 4-year-old containing initial /p/. Use words that can easily be represented with pictures or toys. Exclude multisyllabic or uncommon vocabulary. Group the words into food, animals, toys, and everyday objects. Do not provide articulation cues or decide whether /p/ is an appropriate therapy target.”
This separates the tasks appropriately.
The SLP selects /p/.
ChatGPT organizes usable stimuli.
Activity Example: Pack the Picnic
Suppose /p/ has already been selected as an appropriate target.
ChatGPT can generate familiar picnic-related words such as:
peach
pear
plate
pizza
peas
pickle
The clinician can place pictures or objects around the room and have the child “pack” each item after producing the selected target.
The treatment value is originated from the clinician's:
target selection;
cueing;
feedback;
number of trials;
and progression.
The picnic theme simply makes the repetitions easier to organize.
Do Use AI to Invent Speech-Sound Norms
A useful prompt should also prevent the model from improvising developmental information.
Rather than asking:
“Which sounds should a 3-year-old have?”
a clinician who needs normative information should consult an authoritative source first.
Large cross-linguistic reviews show that consonant acquisition is variable across languages and that simple internet charts can overstate developmental acquisition.
Current ASHA guidance treats developmental order as one source of clinical information rather than a diagnostic threshold.
AI should only work based on clinician-supplied parameters.
Better Prompt
“I have already selected /m, b, p/ as treatment targets based on my assessment. Generate three play activities that each allow at least 20 opportunities to produce words containing these targets. Keep vocabulary appropriate for preschool children.”
That is a much stronger use of generative AI.
Do Not Use Bubbles as Speech Treatment Simply Because the Target Is /b/ or /p/
Blowing bubbles is sometimes used as an example of practicing /b/.
Bubbles can certainly make a therapy activity fun.
But blowing bubbles is not the same motor task as producing /b/.
If the target is speech production, the activity should create opportunities for actual speech.
For example:
“Pop the big blue bubble.”
now contains multiple speech targets.
The bubble is the activity.
The speech productions are the treatment.
This prevents an AI-generated game from drifting into unsupported nonspeech exercises.
2. The Bottom-Up or Discrete-Skill Perspective
A bottom-up perspective assumes that a complex behavior can be decomposed into smaller component skills that are established and then integrated into increasingly complex performance.
For an articulatory target, that logic may look familiar:
establish sound → syllable → word → phrase → sentence → conversation
ASHA describes a similar broad progression for many speech-sound treatments:
establishment → generalization → maintenance.
This perspective is particularly intuitive when a child has a specific motor-based distortion and needs to establish an accurate production before using it automatically.
But it should not be confused with every structured speech-sound treatment.
Minimal Pairs Are Not Articulation Drills
Minimal pairs are sometimes placed inside the bottom-up approach and paired with isolated practice of speech sounds.
That combines two different treatment logics.
Minimal-pair therapy is a phonological contrast approach.
The objective is not to perfect the sound as an isolated motor movement. The objective is to learn that the sound contrast changes the word meaning.
tea — key
tar — car
tap — cap
ASHA identifies minimal oppositions as one of several phonological-contrast approaches. Current guidance suggests conventional minimal pairs are particularly appropriate for children with a relatively small number of speech errors.
So ChatGPT can generate minimal pairs — but only after the clinician has determined that minimal-pair treatment actually fits the child's phonological profile.
Better Prompt for Minimal Pairs
Generative models can create plausible-looking word pairs that are poor treatment stimuli.
A better prompt constrains the output tightly.
“Generate 15 real-word English minimal pairs contrasting word-initial /t/ and /k/. Each pair must differ only in the initial target phoneme. Use common words that can be pictured for a preschool or early-elementary child. Verify that both items are real words. Do not include near-minimal pairs. Present the IPA transcription beside each pair so I can independently check the contrast.”
The final sentence supports that the SLP still verifies the list.
AI-generated phonetic material should not be accepted simply because it looks correct.
Example: Minimal-Pair Communication Game
Once the clinician approves the targets, the activity can demonstrate the communicative function of the contrast.
Place pictures of:
tea / key
tar / car
tap / cap
in front of the child.
The child tells the clinician which item they want.
If key is produced as tea, the listener's response can highlight the ambiguity.
This follows the logic of phonological-contrast intervention: speech sounds distinguish words and therefore distinguish meaning.
Current research distinguish conventional minimal pairs from maximal and multiple oppositions, which may be more appropriate when the child has broader or more severe phonological errors.
Better Prompt for Articulation Hierarchy
For a child with a motor-based single-sound target, a different prompt is appropriate:
“The treatment target is initial /ʃ/. The child can already produce /ʃ/ accurately in single words with moderate cueing. Generate 15 short phrases containing one initial /ʃ/ target each, followed by 10 simple self-generated sentence starters. Keep vocabulary appropriate for a 7-year-old. Do not provide tongue-placement instructions.”
This respects the SLP's clinical role.
The clinician establishes the production and determines the cue.
AI generates increasingly complex linguistic contexts.
3. The Language-Based Perspective
A language-based perspective treats speech as part of the broader language system rather than as a collection of isolated motor productions.
This can be particularly useful when the clinician wants to integrate speech treatment with:
phonological contrasts;
phonological awareness;
vocabulary;
narratives;
word structure;
or functional communication.
But simply adding /s/ words inside a story does not automatically make an intervention language based. The clinician should identify which language area is being targeted along with speech.
Interactive Stories Can Support Generalization
Stories become particularly useful once a child can produce the speech target with some accuracy.
Suppose the child is working on /s/ across words and phrases. AI can create a passage with a controlled concentration of /s/ targets.
Better Prompt:
“Write a 150-word story for a 6-year-old containing frequent /s/ targets. Include approximately equal numbers of /s/ in initial, medial, and final word positions. Avoid /s/ clusters because those have not yet been established. Use familiar vocabulary. Bold each target word. After the story, add five retell questions that encourage the child to reuse the target words naturally.”
That is considerably more clinically useful than:
“Write a fun story about a snake with a bunch of /s/ sound words.”
The prompt specifies the phonetic distribution and current treatment boundary.
AI Can Make Target Density Easier to Control
One genuine advantage of generative AI is the ability to request unusually specific stimulus characteristics.
For example:
“Generate a passage containing 20 prevocalic /r/ targets but no postvocalic /r/.”
Or:
“Write 12 sentences containing final /k/ without /k/ clusters.”
Or:
“Create ten questions whose likely answers contain initial /l/.”
This can save preparation time. The clinician still needs to inspect the output.
A model may:
misclassify word position;
include an unintended target;
select an unfavorable phonetic context;
use vocabulary inappropriate for the child;
or generate a word the child cannot understand.
AI increases the speed of stimulus production.
It does not remove quality control.
Speech Practice Can Be Integrated With Narrative Language
For some children, an SLP may want to address both speech production and broader language.
Suppose the child is working on /s/ sound and narrative organization.
A prompt might ask:
“Create six picture-scene descriptions forming a simple story with a character, setting, problem, three events, and solution. Each scene should naturally elicit two or three words containing initial or final /s/. Do not force awkward /s/ vocabulary.”
Now the clinician can collect speech data while also targeting:
sequencing;
story grammar;
causal language;
and retell.
This is more sophisticated than inserting target words indiscriminately into a story.
Integrating Phonological Awareness
Speech-sound disorders can also co-occur with phonological awareness and literacy.
ASHA includes integrated phonological awareness among current SSD treatment options and recommends considering literacy-related abilities when appropriate.
ChatGPT can create materials such as:
rhyme sets;
syllable segmentation activities;
initial-sound comparisons;
phoneme blending tasks;
and letter-sound activities.
Better Prompt
“Create a phonological-awareness activity for a 5-year-old targeting initial-sound matching. Use only familiar one- or two-syllable words. This activity is for phonological awareness rather than articulation.”
This will help the child recognize that sun and sock begin with the same phoneme even if their own /s/ production is distorted.
4. The Broad-Based Perspective
The broad-based perspective is sometimes interpreted as: “Use a little bit of everything.”
That is too loose.
A broad-based treatment plan may address several interacting contributors to communication performance rather than assuming that correcting an individual sound automatically solves the child's broader difficulties.
Depending on the child, this might include:
speech production;
speech perception;
phonological organization;
phonological awareness;
intelligibility;
language;
and communicative participation.
Broad-based treatment therefore needs a clinical rationale tying the components together.
ChatGPT Can Build One Activity Around Several Clinician-Selected Goals
Let's say an SLP has assessed a student and deliberately selected three targets: AI can help design one activity incorporating all three.
Prompt:
“Design a 15-minute supermarket role-play for a 7-year-old. The established therapy targets are /sp/, /st/, and /sk/ clusters. Include 20 opportunities to say words containing these clusters. Add five phonological-awareness questions asking the child to identify the first two sounds in a word. Include three moments when the child rates their own speech as clear or unclear. Do not introduce additional speech targets.”
The SLP has determined the treatment plan.
ChatGPT coordinates the materials.
Broad-Based Does Not Mean Maximizing Cognitive Load
It might be too much asking AI to create a game that targets:
articulation;
grammar;
vocabulary;
phonological awareness;
social skills;
memory;
and storytelling
all at once.
The resulting activity may look comprehensive while producing very little focused practice.
Speech-sound learning often requires substantial numbers of opportunities.
If a 20-minute game contains only six actual speech productions because the rest of the time is spent answering vocabulary questions and following game rules, the activity may be engaging but inefficient.
Current SSD literature emphasizes that intervention decisions include not only target selection but also the structure, practice amount, and organization of treatment.
AI-generated variety should not reduce the dose of the target the clinician intended to teach.
5. The Complexity-Based Perspective
The complexity approach is one of the most interesting—and most frequently oversimplified—SSD treatment frameworks.
Traditional developmental logic might suggest teaching easier or earlier-acquired sounds first.
The complexity approach deliberately considers whether treatment of more complex or less established phonological structures can produce broader generalization to untreated, simpler structures.
Research on complexity has repeatedly demonstrated that treating carefully selected complex targets can trigger systemwide phonological change.
Complexity is not synonymous with “pick a hard sound.” Using /r/ as a complex target simply because it is typically acquired later is not enough.
Storkel's clinical tutorial on the complexity approach identifies several relevant target characteristics, including:
later acquisition;
linguistic markedness;
low existing knowledge or accuracy;
and low stimulability.
Clusters can also carry important complexity relationships.
Therefore, an SLP does not apply complexity treatment simply by looking at a developmental chart and choosing whichever sound appears last. The clinician has to analyze the child's phonological system and predicted generalization relationships.
Complexity Is About Generalization
The reason for selecting a complex target is not to make therapy harder.
It is to attempt to produce broader change.
Gierut's complexity research showed that treatment of more complex structures can generalize to untreated structures that are less complex, whereas the reverse pattern occurs much less consistently.
That makes complexity conceptually different from a developmental progression.
A developmental approach may say: “Teach what normally comes earlier.”
A complexity approach may say: “Teach the target predicted to reorganize the greatest part of the child's speech sound system.”
Those are different clinical hypotheses.
ChatGPT Should Not Select the Complex Target
This is one of the clearest boundaries for AI use.
Do not ask: “Which complex sound should I target for this child?”
unless the output is being treated as brainstorming that will be independently analyzed—not a clinical recommendation.
The model does not have access to a reliable phonological analysis unless the clinician provides one, and even then clinical interpretation remains the SLP's responsibility.
A much safer use is:
“I have selected /sl/ clusters based on a complexity analysis. Generate 25 familiar English words containing word-initial /sl/. Separate one-syllable from multisyllabic words. Provide IPA transcriptions for verification. Do not substitute other clusters.”
Now AI performs a constrained language-generation task.
Prompt for Complexity-Based Generalization
Once the clinician selects the target:
“The treatment target is word-initial /sp/ as part of a complexity-based phonological intervention. Generate 15 real-word targets suitable for a 5-year-old and 10 short phrases containing those targets. Do not introduce /st/ or /sk/ because I need to probe those separately as untreated generalization targets.”
This is a sophisticated use of AI because the prompt protects the experimental logic of therapy.
If the clinician intends to measure whether /sp/ treatment generalizes to /st/, the AI-generated practice material should not accidentally train /st/.
Why AI Is Useful for Probe Design
Generative AI can help clinicians rapidly draft:
treated word sets;
untreated probe words;
generalization passages;
and novel contexts.
Probe construction requires careful clinician verification.
If a clinician is measuring generalization, a supposedly untreated probe cannot accidentally include practiced words or structures.
A useful prompt might be:
“Generate 15 untreated probe words for initial /st/ and /sk/. Exclude every word in this treatment list: [paste de-identified target list]. Keep the probe words age-appropriate and do not create nonwords.”
The SLP then checks the output before use.
AI Prompts Should Support Clinical Decisions
The difference between poor and good prompting can be summarized simply.
Weak:
“Give me speech therapy for a child who can't say /k/.”
Stronger:
“I have selected initial /k/ at the word level as the current treatment target after assessment. The child produces /k/ with approximately 70% accuracy in structured single words with one verbal cue. Generate 20 familiar initial-/k/ words and two activities that allow at least 30 production opportunities. Do not change the cueing strategy or introduce other treatment targets.”
The stronger prompt provides:
the clinician-selected target;
current treatment level;
baseline;
desired material;
practice amount;
and boundaries.
ChatGPT does not have to infer the treatment plan.
An AI Prompt Formula for SLPs
For speech-sound material generation, a useful prompt often contains seven pieces of information.
1. Define the Role of AI
“Generate treatment materials; do not make diagnostic or treatment-selection recommendations.”
2. Give the Clinician-Selected Target
“The target is initial /ʃ/.”
3. State the Current Level
“The child is practicing at the phrase level.”
4. Set Linguistic Constraints
“Use familiar vocabulary for a 6-year-old.”
5. Set Phonetic Constraints
“Avoid /ʃ/ clusters and avoid words containing /r/.”
6. Specify the Required Dose
“Create enough items for at least 40 target productions.”
7. Ask for Verification-Friendly Output
“Bold target words and place them in a separate checklist so I can review them before the session.”
This produces far more clinically useful material than asking for a “fun speech game.”
AI Can Be Confidently Wrong
In 2025, Ramachandra and colleagues tested ChatGPT-3.5 on 162 questions about communication sciences and disorders that were created and rated by subject-matter experts.
The median ratings were relatively high.
But only 45.68% of responses were judged fully correct, and only 54.32% were rated fully comprehensive.
The authors concluded that clinicians should not depend on ChatGPT for accurate and complete answers to complex clinical questions and cautioned that inexperienced clinicians may have particular difficulty identifying incorrect outputs.
The study used an earlier ChatGPT model and therefore does not establish current-model performance.
Its central lesson nevertheless remains relevant:
fluency is not verification.
A confident explanation can still contain a subtle clinical error.
Low-Risk AI Use vs. High-Risk AI Use
Not every AI task carries the same risk.
Lower-Risk Uses
Generating:
word lists;
stories;
conversation prompts;
picture descriptions;
game themes;
home-practice ideas based on clinician-selected targets;
sentence stimuli;
role-play scenarios;
or alternative versions of existing materials.
Higher-Risk Uses
Asking AI to:
diagnose a disorder;
interpret a standardized score;
choose a phonological treatment approach;
determine medical necessity;
decide whether a child has apraxia;
identify the correct target based only on age;
recommend discharge;
or replace the clinician's evidence appraisal.
ASHA's current generative-AI guidance states that clinicians remain responsible for assessing the efficacy and appropriateness of technology and that AI cannot replace the audiologist or SLP.


.png)
.png)