TeleVRSLUBench is a spoken language understanding benchmark that incorporates visual scene information and explicit reasoning processes for joint intent detection and slot filling. The dataset was proposed by Tele-AI in the paper 'Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding'. It was last updated on the Hugging Face platform on 2026-01-22.
Use Cases
- Training models for joint intent detection and slot filling based on multimodal inputs.
- Benchmarking spoken language understanding systems that incorporate visual scene context.
- Researching explicit reasoning processes in multimodal AI tasks.
- Developing more realistic human-computer dialogue systems that use visual grounding.
Strengths
- Integrates scene-level visual context, which the description states is a first for an SLU benchmark.
- Designed for joint intent detection and slot filling, a core task in spoken language understanding.
- Proposed in a dedicated research paper, suggesting a structured academic foundation.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- Tele-AI
- Collection Method
- Proposed in a research paper; specific collection method not detailed in the provided description.
- Freshness
- Last updated 2026-01-22 07:07:25; freshness should be verified.