On air Voice data foundry · for labs

The voice data your model is starving for.

Born-digital, dual-channel conversational audio in the accents Big Tech hasn't covered. Pidgin, Yoruba, Swahili, Zulu and more, verified by native-speaker editors, zero PII by design, ready for your training pipeline.

WAV + JSONL + Parquet · HF compatible GDPR · NDPA · POPIA aligned Consent cryptographically logged
zero PII · consent signed
24 kHz · dual channel · editor verified
Live corpus acc_sess_20260828_pcm_00892.jsonl
{ "locale": "pcm-NG", "duration_seconds": 184.5, "channels": 2, "sample_rate_hz": 24000, "turns": [ { "speaker": "customer", "emotion": "frustrated", "verified_text": "Good morning. I wake up see say POS commot 50k for my account!", "english_source": "I woke up and saw ₦50,000 leave my account…", "inter_turn_latency_ms": 350 }, // 11 more turns, all aligned + tagged ] }
01
01 · The gap

Voice AI has an accent problem.
We fix the supply.

The world's most spoken languages are also the most under-served in training data. Voice-to-voice models fumble the moment they meet a Nigerian customer or a Nairobi merchant. That gap is a supply problem, and it's what we build.

80%+
of open voice-AI training hours are English or Mandarin

Everyone builds voice agents. Almost no one builds them for accents outside Silicon Valley. Legacy call-center data is dirty and PII-laden. Common Voice is single-utterance and monotone. What labs need, real conversational, dual-channel, dialectal audio, barely exists at scale.

We generate it born-digital: gamified crowd, native-speaker verification, word-level alignment on tap.

Open voice-AI training hours · sampled corpora
English
~30k hrs
Mandarin
~18k hrs
Spanish
~9k hrs
Arabic
~5k hrs
Swahili
~1.5k hrs
Yoruba
~800 hrs
Pidgin (NG)
~400 hrs
Zulu
~200 hrs
Aggregated across Common Voice, VoxPopuli, MLS, LibriSpeech and public corpora as of 2026. Non-Western dialects rounded to nearest 100 hrs.
02
02 · The impact

The billions your AI cannot talk to yet.

Voice AI trained only on English and Mandarin structurally excludes most of the world. Local-language voice data doesn't just improve accuracy. It defines whether entire populations can use your product at all.

739M

Adults globally with limited literacy. 77% live in Sub-Saharan Africa and Central/Southern Asia. Voice is their only route into AI.

2,000+

African languages, most primarily spoken, not written. Text-first AI excludes them by design.

Older adults are twice as likely to use voice assistants (51%) as text chatbots. Voice reads as natural, screens as intimidating.

$53B

Global voice AI market projected by 2030. Non-Western coverage is the biggest unclaimed slice.

Voice agent shipping into Nigeria today

Pick your market. Watch the reach change.

Same product, same launch budget. One trained only on English, one trained on the country's spoken languages alongside it. The difference is not a rounding error, it is a different company.

Without local-language voice data

English-only voice agent

  • ~15–25% of Nigerians comfortable transacting in English
  • ~35M addressable customers, out of ~220M population
  • Elderly, rural and low-literacy users silently drop off
  • CSAT sinks on accented English calls, churn hides as "poor fit"
Addressable market
~35M
With Accent Studio corpora

Multi-dialect voice agent

  • ~90%+ of Nigerians served in their preferred spoken language
  • ~200M addressable customers
  • Elderly, rural and low-literacy users fully unlocked
  • Pidgin code-switching handled natively, not filtered out
Addressable market
~200M
5.7×

market expansion in Nigeria alone. Add Nigerian Pidgin, Yoruba, Igbo and Hausa and every excluded speaker becomes a served customer. Layer more countries and the math compounds.

Conservative addressable-market estimates. Sources: UN World Population Prospects 2024 · EF English Proficiency Index 2024 · World Bank adult literacy data · UNESCO Institute for Statistics · WHO Global Report on Health Equity for Persons with Disabilities · Grand View Research, voice AI market

Who your product finally reaches

Four populations voice AI can serve only in their own language.

👵

Elderly

130M+

Africans aged 60+ by 2030. Screen-averse, voice-native, the largest healthcare cost centre.

🌾

Rural & low-literacy

~570M

Adults across Africa and South Asia with limited literacy. Voice is their only interface, in their own language.

Accessibility-first

1.3B

People globally with a significant disability (WHO). Voice-first is a primary access modality.

📱

Mobile-first youth

60%+

Of Africa is under 25. Voice-native from TikTok and WhatsApp, with zero patience for an agent that doesn't get their accent.

03
03 · Hear the gap

What AI thinks it sounds like, versus the real thing.

Same line, two takes. First, a state-of-the-art voice model attempting it. Then a native speaker from our studio. Watch the words light up as they're spoken.

English source promptGood morning. I woke up this morning and saw that ₦50,000 left my account through a POS I never used.

Tap a play button. Words light up as they are spoken.
04
04 · The pipeline

Two players, one linguist, one clean take.

Every session is a small production. Two players improvise from a card in their own language, a trained linguist verifies the transcript, and an audit layer samples everything behind them.

01

Live improv, dual track

Two anonymously paired native speakers play a scene from a card authored in the target language by a linguist, presented as audio with text as a fallback: a fraud dispute, a market haggle, a KYC call. Nothing is translated from English, so the speech is native, not calqued. Each voice records to its own channel.

output · raw dual-channel audio
02

A linguist verifies every transcript

Editors are recruited, trained and certified language professionals, not players. Each transcript is verified by one of them, with an English gloss written downstream from what was actually said, and slang, code-switches and tone tags corrected. Peer ratings below 4 force a retake before anything ships.

output · verified transcript + labels
03

Packaged for training

Sessions are chunked, indexed and hashed, with word-level alignment as a tier option. Every session manifest records its prompt_direction, target-native by default or English-anchored on request. A 30% random audit re-reviews turns blind. Anything that fails is reworked or dropped, never shipped.

output · WAV + JSONL + Parquet bundle
05
05 · The deliverable

Clean, labelled, licensable.
Slots into your pipeline on day one.

Turn schema manifest.jsonl
{ "turn_id": 4, "speaker_id": "spk_pcm_ng_99104_m", "channel": 1, "start_ms": 13700, "end_ms": 19940, "raw_stt_text": "sir i see three unauthorized…", "verified_text": "Oga, I see three unauthorized transactions. I go block am now.", "english_source": "Sir, I can see three unauthorized transactions…", "emotion_label": "focused_professional", "inter_turn_latency_ms": 480, "code_switches": ["EN→PCM", "PCM→EN"], "peer_rating": { "aggregate": 5.0 }, "alignments": [ /* word level, typed */ ] }
AudioDual-channel WAV · 24 kHz · PCM 16-bit · speaker per channel
MetadataJSONL, one row per session · full turn objects
IndexParquet, turn level, dictionary encoded for filter pushdown
LabelsEmotion (10 class) · latency · code-switch spans · peer QA
ConsentPer-speaker signed log, SHA-256 hash anchored
IntegritySHA-256 sums + PGP signature over the bundle
LicencePseudonymous speakers · no re-identification · no cloning of individuals · biometric-grade handling
DeliveryPresigned S3 · TLS 1.3 · clean rooms on request
ToolingHugging Face Datasets, PyTorch and NeMo recipes included

See a full delivery before you spend a dollar.

A reference delivery bundle showing exactly what a buyer receives, with synthetic sessions: 10 sessions, 118 turns across 6 locales, schema reference, indicative evaluation baselines and the consent register. Audio ships on request.

PDF · 46 pages · 2.7 MB · no email required

06
06 · Aligned corpus tier

Word-level alignment, source to target.

Every translated span links back to its English source, typed by phenomenon. Serial verbs, aspect shifts, null particles, the constructions that make these languages hard, mapped for you.

English source
Good morning.I woke upand saw that₦50,000leftmy account
Pidgin · verified take
Good morning.I wake upsee sayPOS commot 50kformy account!
Matching colors are aligned spans. "see say" is a serial-verb construction, "I wake up" an aspect shift, and the dashed "for" a Pidgin particle with no English source. 1,800+ typed alignment spans ship in the sample bundle alone, each with an aligner confidence score.
07
07 · Coverage

Locales and domains we record in.

Locales · live and ramping

Nigerian Pidgin Yoruba Igbo Hausa Swahili · KE Swahili · TZ Zulu Xhosa Amharic Twi Nigerian English Kenyan EnglishBrazilian Portuguese · soonTagalog · soonBahasa · soon

Domains · scenario families

Retail banking · fraudTelco supportFintech KYCMobile moneyMarket haggleRide-hailingInsurance claimsLogisticsConsumer loansHealthcare intakeCustom · commissioned
08
08 · How to buy

Four ways in, no lock-in.

Production is commissioned against a signed order, not pulled from stock. Non-exclusive by default, exclusivity available and priced. Deposit on signature, balance on delivery or in milestones. We quote a lead time and a sustained weekly rate per language from measured throughput.

Start here

Pilot corpus

A scoped bundle in your target locale and domain, sized for a fine-tune experiment. Commissioned on signature; a standing Nigerian Pidgin corpus lets a Pidgin pilot ship within days.

one-off · evaluation friendly
Scale

Standing subscription

A monthly flow of fresh verified hours in the locales you pick. Same schema, versioned releases, replacement SLA.

recurring · volume priced
Bespoke

Custom commission

Your scenarios, personas and edge cases, authored into target-language scenario cards by our linguists. Recruited, interviewed speakers play your test set into existence.

directed · your IP guardrails
Premium

Aligned corpus

Everything above plus word-level alignment, typed by linguistic phenomenon and confidence scored. English-anchored prompting available per order for source-anchored parallel data, recorded per session.

tier 4 · translation grade
Compliance

Consent-cryptographed. Procurement ready.

Consent

Per-speaker signed agreement, SHA-256 anchored, withdrawal SLA

Privacy law

GDPR, NDPA and POPIA aligned, PII redacted at capture

Security

AES-256 at rest, TLS 1.3 in transit, IP-whitelisted delivery

SOC 2

Type II program in progress with a Big Four auditor

Provenance

Born-digital, zero scraped audio, full chain of custody

Read what our contributors sign

The Contributor Agreement, in full.

Every player signs this before their first paid recording: a plain-language summary, explicit biometric consent as a separate act, what licensees may never do, withdrawal, retention, transfers, and the consent record whose fingerprint travels inside every session you receive.

Download the agreement
Version1.1 · effective 11 September 2026 · 11 pages · document ASC-CA-1.1 · supersedes v1.0, which no contributor signed
Verify your copySHA-256 bb776c828a00da7a594d40357df6aa5e615ef1eaaef6926fc1a8273fba297b15
Drafted againstGDPR and UK GDPR · Nigeria Data Protection Act 2023 · Kenya Data Protection Act 2019 · South Africa POPIA · Illinois BIPA. Clause-by-clause map in Annex B.
Consent quality · self-assessed
7of 7

The EDPB Guidelines 05/2020 criteria for valid consent, each with the clause that satisfies it.

  • Freely givencl. 7, 8, 15
    Withdrawal without detriment; verified earnings are paid even when another session is questioned.
  • Specificcl. 3
    Purposes are a closed list. No advertising, profiling or automated decisions.
  • Informedsummary, cl. 1 to 15
    Plain-language summary first, then the full terms, including who hears the recording and how pay is set.
  • Unambiguoussignature block
    Two affirmative ticks plus a one-time code to the registered email.
  • Explicit, for biometric datacl. 4
    A separate consent act, with the legal bases named.
  • As easy to withdraw as to givecl. 8
    One tap in the app or one email. Undelivered audio deleted in 30 days.
  • DemonstrableAnnex A
    Append-only consent record; its SHA-256 is written into every session manifest.

Self-assessment by Accent Studio against a published standard, v1.0. External attestations will be listed here as they complete. Not legal advice.

The details

What procurement and ML leads ask us.

What formats do you deliver?
Dual-channel WAV at 24 kHz PCM 16-bit, JSONL manifests with full turn objects, a Parquet index for filtering and streaming, plus consent logs, checksums and a schema reference. Hugging Face Datasets, PyTorch DataLoader and NeMo recipes ship in the bundle.
How do you handle PII?
Scenes are synthetic scenarios, so real PII should never appear. Anything that slips through, a real phone number, an ID, is caught at capture and replaced with synthetic tokens before storage. Speakers are indexed by cryptographic UUID only, and the identity vault is never delivered.
Can we specify custom scenarios or personas?
Yes. Custom commissions drop your scenarios into the game as scenario cards. You define the domain, personas, twists and target register, and the crowd plays it into existence with the same verification pipeline on top.
How is data delivered?
Presigned S3 URLs over TLS 1.3 with IP whitelisting by default. AWS Clean Rooms, Snowflake Data Clean Rooms, GCS and Azure SAS available. Enterprise buyers can request air-gapped drive shipment.
How fast can we get a pilot corpus?
Weeks, not quarters. The crowd is already playing in our live locales, so a pilot in an existing locale and domain is mostly packaging time. New locales or custom scenario families add ramp time and we'll quote it honestly.
What are the usage rights?
Commercial, non-exclusive, perpetual and worldwide by default: train, fine-tune, evaluate and ship models freely. Redistribution of the raw dataset is prohibited, model outputs are yours. Exclusive tiers are negotiable per corpus.
Do you handle data residency?
Primary storage is us-east-1 with an eu-west-1 replica on request. Region-pinned delivery and processing agreements are available for enterprise contracts.
Talk to us

Tell us what you're building.

We'll come back within one business day with a sample bundle, a schema walkthrough and an honest read on fit.