Speech corpus

ConvoAAVE

A community-built speech corpus of African American Vernacular English, developed to improve how speech recognition systems handle the dialect. Recordings remain private to protect participants; the transcript layer is published.

ConvoAAVE began with a question students raised repeatedly in our programs about why these systems misunderstood them. Published audits give the answer: commercial speech recognition shows measurably higher error rates for Black speakers than for white speakers. Voice assistants, captioning, dictation, and classroom transcription tools all inherit that gap.

The cause is representation. Speech models learn from audio, and the audio used to train them underrepresents AAVE speakers. A system that has rarely encountered habitual be, stressed been, or finna does not treat them as grammar, and normalizes them into forms the speaker did not use.

Working with churches, schools, community centers, and family networks across Hartford and East Hartford, we recorded unscripted conversation and extended narrative speech, then transcribed it as spoken. The result is more than 80,000 words of transcript, paired with systematic coverage of the dialect’s grammatical constructions.

a Black speakers White speakers 0 0.1 0.2 0.3 0.4 0.5 Word error rate 0.35 0.19 Mean, five systems 0.45 0.23 Apple, worst 0.27 0.15 Microsoft, best mean for black speakers, 0.35 Error rates for black speakers are nearly twice those for white speakers in every system tested. b 0 0.1 0.2 0.3 Word error rate 0.28 0.12 Apple 0.21 0.10 IBM 0.17 0.11 Google 0.18 0.08 Amazon 0.13 0.07 Microsoft Identical short phrases, spoken by both groups. The gap is not what people said.
Fig. 1 Commercial speech recognition is close to twice as wrong for black speakers as for white speakers. a, Average word error rate over a matched sample of 2,141 audio snippets per group, drawn from 19.8 hours of interview audio with 73 black and 42 white speakers. Apple performed worst overall and Microsoft best, so the three systems not plotted fall between them. The largest standard error among the ten values published is 0.005. b, The same measurement restricted to identical short phrases spoken by both groups, which rules out what people said as the explanation. Data from Koenecke, A. et al. Racial disparities in automated speech recognition. PNAS 117(14), 2020. doi:10.1073/pnas.1915768117
a Black speakers 23 in 100 unusable White speakers 1.6 in 100 unusable Unusable means a word error rate above 0.5, the point past which a transcript cannot be relied on. Each circle is one audio snippet in a hundred. b Black speakers White speakers 0 0.1 0.2 0.3 0.4 0.5 Word error rate 0.41 0.21 Men 0.30 0.17 Women The gap is widest for black men, the speakers these systems transcribe worst.
Fig. 2 The failure is not spread evenly, and for close to one recording in four it is total. a, Share of snippets whose transcript is unusable, taking a word error rate above 0.5 as the point past which a transcript cannot be relied on. Ten times as many snippets of black speakers fail that bar. b, Average error rate by gender, over the same five systems. Data from Koenecke, A. et al. Racial disparities in automated speech recognition. PNAS 117(14), 2020. doi:10.1073/pnas.1915768117

How we built it

  1. Community partnershipsWe met with churches, schools, community centers, and family networks to explain the project and invite participation, rather than recording people who had not agreed to it.

  2. Consent and privacy protocolParticipation was voluntary and revocable. Every speaker was told what the recording was for, and who would be able to hear it, before anything was captured.

  3. Recording natural speechWe recorded real conversation and extended storytelling instead of read prompts, so the corpus carries the rhythm, overlap, and structure of how people actually speak.

  4. Transcription as spokenRecordings were transcribed to preserve AAVE grammar rather than normalize it into Standard English, since a transcript that corrects the speaker removes the signal the corpus exists to provide.

  5. Construction coverageAlongside the conversations we assembled systematic examples of AAVE constructions, including habitual be, stressed been, completive done, finna, stay, negative concord, and might could.

  6. Evaluation and return to the communityWe use the corpus to test how well systems handle the dialect, and bring the results back to the classrooms and families who built it.

Why the audio is not publicVoice recordings identify the speaker, and these were made by community members discussing their own lives. Publishing the audio would expose participants in order to make the corpus more convenient to use. We release the transcripts, document the methodology, and keep the recordings under restricted access.

The seven constructions the corpus documents
Construction As spoken As a normalizing system returns it What the normalization costs
1 Habitual be she be working weekends she is working weekends the habit is lost: one shift, not every weekend
2 Stressed been I been knowing her I have known her the remote past is flattened into the recent past
3 Completive done he done finished it he finished it completion, and the speaker’s stance on it, drop out
4 Finna we finna leave we are going to leave immediacy is lost, and the token is often misheard outright
5 Stay they stay busy they are busy a persistent state is read as a passing one
6 Negative concord he don’t know nothing he does not know anything agreement is rewritten as an error rather than read as grammar
7 Might could I might could come I could come the double modal, and the hedge it carries, disappears
Fig. 3 Seven constructions the corpus documents, and the grammatical information a normalizing transcript throws away. Each row is a construction the annotation layer covers. The middle columns show the form as a speaker produces it and the form a system that has not learned the construction returns in its place. The examples are constructed to illustrate each pattern; they are not excerpts from the corpus, which is why transcription as spoken is treated here as a measurement decision rather than a style choice. A system scored on a normalized transcript is scored against a sentence nobody said.
The corpus

Eighty thousand words, and seven ways in

Drag to turn

Every dot is one word of the transcript we hold: more than 80,000 of them. The seven markers are the constructions the annotation layer tags. Take one and it pulls its own thread of words out of the body, which is what the annotation layer is for.

Choose a construction to pull its thread out of the transcript.

The dots stand for a count we publish. Their positions do not: we do not publish how many words carry each construction, so the seven threads are drawn the same size rather than implying a distribution we have not measured. No recording or transcript line is reproduced here.

Method

How the corpus is built

The design is the argument. What gets recorded, what gets released, and how the transcript is written all decide whether the corpus can measure the thing it exists to measure.

a Community churches, schools, centers, families 1 Consent voluntary and revocable 2 Recording unscripted talk, not read prompts 3 Transcription as spoken, not normalized 4 Annotation seven AAVE constructions 5 Evaluation tested, then returned 6 b AUDIO restricted access, never published a recorded voice identifies its speaker stops here TRANSCRIPT more than 80,000 words, published de-identified, AAVE grammar preserved RECORDED IN Hartford and East Hartford, with written consent, in churches, schools, community centers, and family networks. Consent is revocable at any stage, and a withdrawal removes the speaker from both the audio and the transcript.
Fig. 4 How the corpus is built, and the point at which the audio stops. a, The six stages, in order. Consent precedes recording, and no session begins with a speaker who has not agreed to it. b, A recording yields two artifacts and only one of them leaves the room. The transcript is de-identified and published, more than 80,000 words of it; the audio is held under restricted access because a recorded voice identifies the person who produced it. Recording took place in Hartford and East Hartford, with written consent.