Labels tied to model behaviour.
- verbatim transcript
- normalized transcript
- language and code-switch tags
- speaker cohort metadata
- device and environment
- noise and overlap events
- utterance intent where applicable
A multilingual speech collection concept spanning Indian English, Hindi and regional-language cohorts across devices and acoustic environments.
Every specification is adapted to the buyer's model, production environment, data rights and measurable acceptance criteria.
Targets are agreed during scoping and reported only after measurement on the applicable delivery.
Buyer-specific schema and storage requirements can be incorporated before production.
audio/manifest.jsonlspeaker_metadata.csvlanguage_map.csvdataset_card.mdqa_benchmark.jsonlShare the model objective, target environment, approximate volume and delivery constraints. We will respond with the questions needed to scope a credible dataset.
Request this dataset pattern ↗