Fluency model
The core model: detects every disfluent moment, then sorts it into the clinical types.
Practice that meets your voice where it is
Evidence-based tools for anyone working on how they speak, whether you stutter, are navigating a speech difference, or simply want to be heard more clearly.
The problem
To score a single sample, a clinician replays the audio, counts every disfluency, sorts each one by type, times the long ones, and works out the percentages by hand. It is careful, skilled work. It is also the reason objective measurement often gets skipped, and the reason progress between sessions goes unrecorded.
Time spent tallying is time not spent on therapy, planning, or the client in the room.
Because hand-counting is slow, assessment rests on short clips rather than how someone speaks all week.
Practice at home and work in the clinic rarely meet in the same set of numbers.
What's inside
The core model: detects every disfluent moment, then sorts it into the clinical types.
The model flags disfluencies on real audio; clinicians confirm the calls and keep the record.
One big record button in your pocket, for individuals at home and clinicians mid-session.
Pacing and disfluency cues as you speak, so a session can be adjusted while it is happening.
Short guided sessions with real-time feedback, built around techniques clinicians already use.
Telehealth sessions stream straight into the Workbench, with live pacing cues during the call.
School, IEP, and progress summaries assembled from confirmed labels, ready to send.
Encrypted in transit and at rest, anonymized before training, and never sold to anyone.
Pacing, clarity, and other speech differences, trained on the confirmed labels built along the way.
For clinicians
Open a recording and the events are already marked: every repetition, prolongation, and block, tagged with a type and a timestamp. Drag a flag to move it, pull its edges to change its span, retype it in one click. The counts and percentages update as you work.
Model-tagged repetitions, prolongations, and blocks, editable by hand.
Stutter-like and total disfluency rates, type distribution, speaking rate, weighted severity.
Confirmed labels become progress summaries, school reports, and session notes.
For individuals
The phone app is a recorder first: one big button, wherever you are. Order a coffee, run a standup, read a passage. The model listens for the same events a clinician would mark, so what happens outside the clinic finally counts.
For developers
Teletherapy platforms, EHRs, research groups, and consumer speech apps can send audio and get structured disfluency events back. Every key chooses its own data path, and healthcare workloads stay sealed by default.
POST /v1/analyze { "audio": "<base64>", "return": ["events", "measures"] } → { "events": [ { "t": 4.12, "type": "block", "dur": 1.4, "conf": 0.86 }, { "t": 7.90, "type": "sound_rep", "dur": 0.6, "conf": 0.74 } ], "measures": { "sld_pct": 4.5, "speech_rate_wpm": 122 } }
How it works
Detection and classification are separate problems, so we treat them separately. Splitting them keeps the hard judgement calls where they belong: with the clinician.
The first model answers one question across the whole recording: is this stretch of speech disfluent or not? A narrow question is a learnable one, and it runs on far more audio than any labelled corpus of types.
Only the detected moments go to the second model, which sorts each one into the clinical categories: sound and word repetitions, prolongations, blocks, and the non-stuttering-like disfluencies alongside them.
No label is final until a qualified SLP accepts it. Confirmations and corrections both feed back as training signal, so the corpus grows as professionally labelled data, not model output grading itself.
Every confirmed label is attributable to a licensed professional. That is what makes the resulting dataset worth training on, and what makes the output defensible in a clinical record.
Research and evidence
The Workbench computes the measures clinicians are trained on, using the standard disfluency taxonomy rather than a proprietary score: stutter-like and total disfluency rates, the ratio between them, prolongations and blocks as a share of stuttering-like events, speaking rate against age norms, and a weighted severity that maps to the usual bands.
Validation is the work in front of us, not behind it. Nothing here is peer-reviewed yet, and we will not pretend otherwise. What we can share now is the methodology, and what we want is clinicians and researchers to help design the study.
The clinical whitepaper is not written yet
Working on validation before the paper.
PlannedPrivacy
Voice is identifying, and clinical voice is protected health information. We are early and we will be plain about it: to start, all model inputs are used as training data. Nothing is sold, everything is anonymized, and the separated paths below are what we are building toward.
Client audio and confirmed labels are handled to HIPAA standards and never sold or shared with other clinics. While we are early, they also train the models; clinic-held sealed paths are on the roadmap.
Practice audio is anonymized and used to improve the models. Per-product opt-outs and separated training paths land as the platform matures.
Audio in, structured events out. While we are early, API audio also trains the models; per-key sealed paths are on the roadmap.
Market and business model
Roughly one in a hundred adults stutters, and many more live with other speech differences. They are served by speech-language pathologists in private practice, schools, hospitals, and increasingly over video, all of whom document outcomes and all of whom do it by hand today.
Subscription per clinician, sold into private practices and school districts. The Workbench pays for itself in review time.
Land with one SLP, expand across the practice
A consumer subscription that also, with consent, produces the practice audio the models learn from.
Distribution and data in the same motion
Usage-based pricing for teletherapy platforms, EHR vendors, and research groups, with a free tier to build against.
Reach clinicians we never sell to directly
The compounding loop: individuals generate speech, clinicians confirm the labels on it, and confirmed labels make the models better for both. Every product makes the next one cheaper to run and harder to copy.
Roadmap
The Workbench, the individual dashboard, the API, and the iOS recorder: model-tagged review with editable flags and the standard measures computed from confirmed labels.
Entity, BAA, and the HIPAA posture behind it, plus the data agreements that let clinics put real client audio through the Workbench.
Design the study with clinicians, measure agreement between model output and clinician labels by disfluency type, then publish the methods and the misses.
The same detect-then-classify architecture applied to other speech differences, on the corpus of professionally confirmed labels built along the way.
Team
StutterStep is early. The product is being built alongside the clinicians who will use it, and the parts that are not ready yet are marked as such on this page.
Founder
Building the models, the Workbench, and the case for both.
Speech and audio ML, and a clinical lead who has counted disfluencies by hand for a living.
Practising SLPs, fluency researchers, and people who stutter, shaping what gets measured and how.
If you assess fluency, teach it, research it, or live it, we want your judgement in the loop early.
Thirty minutes with the Workbench and your own opinions about what it gets wrong.
Pick a time →Universities, clinics, health organisations, and investors: there is a version of this that involves you.
Start a conversation →Shape the taxonomy, the measures, and the line between model and clinician.
Put your name in →© 2026 StutterStep, Inc. stutterstep.ai