A community-built speech corpus of African American Vernacular English, developed to improve how speech recognition systems handle the dialect. Recordings remain private to protect participants; the transcript layer is published.
ConvoAAVE began with a question students raised repeatedly in our programs about why these systems misunderstood them. Published audits give the answer: commercial speech recognition shows measurably higher error rates for Black speakers than for white speakers. Voice assistants, captioning, dictation, and classroom transcription tools all inherit that gap.
The cause is representation. Speech models learn from audio, and the audio used to train them underrepresents AAVE speakers. A system that has rarely encountered habitual be, stressed been, or finna does not treat them as grammar, and normalizes them into forms the speaker did not use.
Working with churches, schools, community centers, and family networks across Hartford and East Hartford, we recorded unscripted conversation and extended narrative speech, then transcribed it as spoken. The result is more than 80,000 words of transcript, paired with systematic coverage of the dialect’s grammatical constructions.
Community partnershipsWe met with churches, schools, community centers, and family networks to explain the project and invite participation, rather than recording people who had not agreed to it.
Consent and privacy protocolParticipation was voluntary and revocable. Every speaker was told what the recording was for, and who would be able to hear it, before anything was captured.
Recording natural speechWe recorded real conversation and extended storytelling instead of read prompts, so the corpus carries the rhythm, overlap, and structure of how people actually speak.
Transcription as spokenRecordings were transcribed to preserve AAVE grammar rather than normalize it into Standard English, since a transcript that corrects the speaker removes the signal the corpus exists to provide.
Construction coverageAlongside the conversations we assembled systematic examples of AAVE constructions, including habitual be, stressed been, completive done, finna, stay, negative concord, and might could.
Evaluation and return to the communityWe use the corpus to test how well systems handle the dialect, and bring the results back to the classrooms and families who built it.
Why the audio is not publicVoice recordings identify the speaker, and these were made by community members discussing their own lives. Publishing the audio would expose participants in order to make the corpus more convenient to use. We release the transcripts, document the methodology, and keep the recordings under restricted access.
| Construction | As spoken | As a normalizing system returns it | What the normalization costs | |
|---|---|---|---|---|
| 1 | Habitual be | she be working weekends | she is working weekends | the habit is lost: one shift, not every weekend |
| 2 | Stressed been | I been knowing her | I have known her | the remote past is flattened into the recent past |
| 3 | Completive done | he done finished it | he finished it | completion, and the speaker’s stance on it, drop out |
| 4 | Finna | we finna leave | we are going to leave | immediacy is lost, and the token is often misheard outright |
| 5 | Stay | they stay busy | they are busy | a persistent state is read as a passing one |
| 6 | Negative concord | he don’t know nothing | he does not know anything | agreement is rewritten as an error rather than read as grammar |
| 7 | Might could | I might could come | I could come | the double modal, and the hedge it carries, disappears |
Drag to turn
Every dot is one word of the transcript we hold: more than 80,000 of them. The seven markers are the constructions the annotation layer tags. Take one and it pulls its own thread of words out of the body, which is what the annotation layer is for.
The dots stand for a count we publish. Their positions do not: we do not publish how many words carry each construction, so the seven threads are drawn the same size rather than implying a distribution we have not measured. No recording or transcript line is reproduced here.
The design is the argument. What gets recorded, what gets released, and how the transcript is written all decide whether the corpus can measure the thing it exists to measure.