Benchmark answer Pick a question. These are real QA pairs from RoadSocial for this exact video.
- Problem
- Understanding a road event takes more than seeing it. It takes knowing where you are, what's normal there, and what nearly happened but didn't.
- Approach
- We convert social commentary on road videos into video question-answer supervision, drawing on contextual insight the footage alone cannot supply.
- Result
- A benchmark that exposes where Video LLMs actually break (timing, reasoning, and hallucination), and points to social narrative as a promising source of supervision for the parts models find hardest.