0.12 px
Median epipolar error of image correspondences against the geometry implied by the released poses. Pairing the same images with a shifted pose gives 2.83 px, so the test is sensitive rather than trivially satisfied.
128,892 real-world clips averaging 78.9 seconds, each released with a camera trajectory and captions grounded at both the clip and the segment level.
1 Alaya Lab · 2 Shanghai Innovation Institute · 3 Wuhan University · 4 Tsinghua University
* Corresponding authors
Scale and composition
The inherited Sekai subset supplies 37.0% of clips but 28.1% of hours, since every clip is exactly 60 s. The newly collected YouTube subset supplies 62.2% of clips and 67.7% of hours. The self-captured 360° subset is 0.8% of clips yet 4.2% of hours, because its captures average 436 s and are released uncut. Perspective clips are anchored at a two-minute cap: the clips that reach it are a third of the corpus by count but 51.4% of all footage.
Open full size ↗
One dataset, many ways to move
Ground walking, driving, rail, cycling, boats, cable cars, escalators and drone flight. The demos below cover each regime, and every clip carries the same pose and caption schema regardless of how the camera moved.
Environment and scene diversity
Every clip is labelled along country, camera motion, lighting, time of day, weather and scene type. The head is dominated — urban streets, clear skies — but the minority values stay populated at scale: 22.4% of clips are shot at night, 41.5% under non-clear weather, 34.0% under artificial or mixed lighting, and 28.7% of the newly collected footage is natural scenery rather than street. Pick an attribute to see its distribution.
· bar length is the share of the 128,892 clips · hover a bar for exact counts
Geographic diversity
113 countries and regions after alias normalization. Japan, the United Kingdom and the United States together supply about two thirds of geographically resolved clips, and 14 countries are needed to reach 90% cumulative coverage — the remaining 110 span every inhabited continent. Click any shaded country for its clip count and footage actually recorded there.
Hover or tap any shaded country · 99 countries are interactive
Explore the collection
Every card is a released clip shown with its own footage. Filter by camera regime, then open one for the full case study: sampled frames, the ViPE trajectory, the clip-level fields and every grounded segment.
Structured semantics
Clip-level fields summarize what persists; 649,597 grounded segments — 5.04 per clip, 15 s median — describe how it changes. The levels are not restatements of one another: among the four factorized descriptions, pairwise vocabulary overlap is only 0.16, and each contributes words no other field uses.
Caption case studies
Each example aligns sampled frames, the released ViPE trajectory, the clip-level fields and the temporally grounded segments, so a change of camera path and a change of description can be read against each other.
Camera as supervision
Because every clip ships a trajectory, motion statistics cover the whole corpus rather than a sample. Newly collected clips travel about twice as far as inherited ones (median path length 123 vs 68 units) and turn more (270° vs 184°), while the 360° captures travel farthest yet rotate least.
Panoramic capture
982 self-captured 360° videos, 119 hours, averaging 436 s. Routes are walked to close: loop closures are geometrically verified and refine 1,051 of 1,283 trajectories (81.9%), which is what makes full-sequence accumulation coherent.
Observations accumulated along the loop-refined poses over a whole capture. Repeated structures align instead of forming separate copies, which is what a small residual drift would otherwise destroy. Color encodes progression along each trajectory.
Validation
Both released modalities are audited on their own terms, and then against each other. Every figure below is inference-only — no model was trained to produce it — and each test is one that could have come out the other way.
Median epipolar error of image correspondences against the geometry implied by the released poses. Pairing the same images with a shifted pose gives 2.83 px, so the test is sensitive rather than trivially satisfied.
Rotation disagreement with an independently executed DROID-SLAM run after Sim(3) alignment, with intrinsics re-estimated for every subset so no source is measured under an easier protocol. Normalized ATE is 0.018.
Of the 74,437 clips carrying both labels, the segments a caption calls turn rotate more than the ones it calls straight — each clip serving as its own control.
Preference over a conventional one-shot caption written by the same model from the same frames, scored in the presentation order that works against us.
| Test | What it establishes | Result | Reference point |
|---|---|---|---|
| Pose validity | No missing or discontinuous tracks, over all 128,892 clips | 1.000 valid ratio | median jump rate 0.00000 |
| Epipolar consistency | Released poses explain the images, without ground truth or a second estimator | 0.12 px median | 2.83 px with a mismatched pose |
| Cross-run agreement | An independent DROID-SLAM run recovers the same geometry under matched intrinsics | 0.103° rotation | ATE 0.018 · 99.5% reconstructed |
| Caption faithfulness | A VLM judge reading 12 frames scores coverage and per-aspect correctness | 4.01–4.31 / 5 | camera behaviour scores lowest |
| Caption preference | The multi-field schema beats a one-shot caption from the same model | 60.9% preferred | order chosen to disfavour us |
| Schema complementarity | The six description views carry different words, not restatements of one another | 0.16 Jaccard overlap | 51% of words appear in one view only |
| Cross-modal grounding | Independently produced captions and poses agree on when the camera turns | 0.709 AUC | yaw 1.3° → 27.9° across labels |
| Within-clip control | The ordering survives when each clip is its own control | 71.8% of clips | median gap 12.7° · 74,437 clips |
| Frame quality | Long footage is not bought with worse images | 5.23 aesthetic | MiraData 5.45 · RE10K 5.20 |
| Frame-to-frame motion | Sustained translation rather than slow orbits around a static scene | 35.1 flow | 1.4–1.8× the prior corpora |
Build with Sekai2
The technical report, dataset access, annotation code and benchmark splits will appear here. Every clip ships video, a ViPE camera trajectory and a hierarchical caption under one schema.