Sekai2: From World Exploration
to Interactive World Modeling

128,892 real-world clips averaging 78.9 seconds, each released with a camera trajectory and captions grounded at both the clip and the segment level.

Kang He1,2,3   Wenshuo Peng4   Zihui Gao1   Jiaming Tan1   Kaipeng Zhang1,*   Yongtao Ge1,*

1 Alaya Lab  ·  2 Shanghai Innovation Institute  ·  3 Wuhan University  ·  4 Tsinghua University
* Corresponding authors

Scroll to explore
128,892video clips
2,826 hreal-world footage
113countries and regions
649,597grounded segments
78.9 smean clip length

Scale and composition

Three sources, ranked
differently by clips and by hours.

The inherited Sekai subset supplies 37.0% of clips but 28.1% of hours, since every clip is exactly 60 s. The newly collected YouTube subset supplies 62.2% of clips and 67.7% of hours. The self-captured 360° subset is 0.8% of clips yet 4.2% of hours, because its captures average 436 s and are released uncut. Perspective clips are anchored at a two-minute cap: the clips that reach it are a third of the corpus by count but 51.4% of all footage.

Sekai2 dataset composition: newly collected YouTube footage 80,211 clips / 1,911 hours spanning 18 motion types across 72 countries, reused Sekai footage 47,699 clips / 795 hours across 107 countries, self-captured Panoramic 360 982 clips / 119 hours, totalling 128,892 clips and 2,826 hours across 113 countries Open full size ↗

One dataset, many ways to move

Real trajectories through
an open, changing world.

Ground walking, driving, rail, cycling, boats, cable cars, escalators and drone flight. The demos below cover each regime, and every clip carries the same pose and caption schema regardless of how the camera moved.

Environment and scene diversity

Six controlled attributes,
long-tailed like the real world.

Every clip is labelled along country, camera motion, lighting, time of day, weather and scene type. The head is dominated — urban streets, clear skies — but the minority values stay populated at scale: 22.4% of clips are shot at night, 41.5% under non-clear weather, 34.0% under artificial or mixed lighting, and 28.7% of the newly collected footage is natural scenery rather than street. Pick an attribute to see its distribution.

· bar length is the share of the 128,892 clips · hover a bar for exact counts

Geographic diversity

Worldwide coverage,
grounded in real scenes.

113 countries and regions after alias normalization. Japan, the United Kingdom and the United States together supply about two thirds of geographically resolved clips, and 14 countries are needed to reach 90% cumulative coverage — the remaining 110 span every inhabited continent. Click any shaded country for its clip count and footage actually recorded there.

fewer clipsmore clips113 countries

Hover or tap any shaded country · 99 countries are interactive

Explore the collection

Many cameras.
Many worlds.

Every card is a released clip shown with its own footage. Filter by camera regime, then open one for the full case study: sampled frames, the ViPE trajectory, the clip-level fields and every grounded segment.

Structured semantics

Describe the world
at two temporal scales.

Clip-level fields summarize what persists; 649,597 grounded segments — 5.04 per clip, 15 s median — describe how it changes. The levels are not restatements of one another: among the four factorized descriptions, pairwise vocabulary overlap is only 0.16, and each contributes words no other field uses.

Example
Global annotation synchronized

Caption case studies

See space, motion,
and language together.

Each example aligns sampled frames, the released ViPE trajectory, the clip-level fields and the temporally grounded segments, so a change of camera path and a change of description can be read against each other.

Camera as supervision

Every clip comes with
a path through 3D space.

Because every clip ships a trajectory, motion statistics cover the whole corpus rather than a sample. Newly collected clips travel about twice as far as inherited ones (median path length 123 vs 68 units) and turn more (270° vs 184°), while the 360° captures travel farthest yet rotate least.

RGB + 3D POSEDrone15 s synchronized window on the complete trajectory
released clip · 1280×720
drag to orbit · scroll to zoom
Start EndColor encodes timeInteractive: orbit · pan · zoom

Panoramic capture

Return to places
you have already seen.

982 self-captured 360° videos, 119 hours, averaging 436 s. Routes are walked to close: loop closures are geometrically verified and refine 1,051 of 1,283 trajectories (81.9%), which is what makes full-sequence accumulation coherent.

native equirectangular · full 360° × 180° field
complete trajectory · synchronized point
360° pose case studies Meander

Panoramic footage preserves the surrounding environment during turns, loops, and repeated visits.

Full-accumulation reconstruction

Long trajectories recover coherent spatial structure.

Observations accumulated along the loop-refined poses over a whole capture. Repeated structures align instead of forming separate copies, which is what a small residual drift would otherwise destroy. Color encodes progression along each trajectory.

Validation

Measured,
not asserted.

Both released modalities are audited on their own terms, and then against each other. Every figure below is inference-only — no model was trained to produce it — and each test is one that could have come out the other way.

01

0.12 px

Median epipolar error of image correspondences against the geometry implied by the released poses. Pairing the same images with a shifted pose gives 2.83 px, so the test is sensitive rather than trivially satisfied.

02

0.103°

Rotation disagreement with an independently executed DROID-SLAM run after Sim(3) alignment, with intrinsics re-estimated for every subset so no source is measured under an easier protocol. Normalized ATE is 0.018.

03

71.8%

Of the 74,437 clips carrying both labels, the segments a caption calls turn rotate more than the ones it calls straight — each clip serving as its own control.

04

60.9%

Preference over a conventional one-shot caption written by the same model from the same frames, scored in the presentation order that works against us.

Every validation result, with what it is measured against
TestWhat it establishesResultReference point
Pose validity No missing or discontinuous tracks, over all 128,892 clips 1.000 valid ratio median jump rate 0.00000
Epipolar consistency Released poses explain the images, without ground truth or a second estimator 0.12 px median 2.83 px with a mismatched pose
Cross-run agreement An independent DROID-SLAM run recovers the same geometry under matched intrinsics 0.103° rotation ATE 0.018 · 99.5% reconstructed
Caption faithfulness A VLM judge reading 12 frames scores coverage and per-aspect correctness 4.01–4.31 / 5 camera behaviour scores lowest
Caption preference The multi-field schema beats a one-shot caption from the same model 60.9% preferred order chosen to disfavour us
Schema complementarity The six description views carry different words, not restatements of one another 0.16 Jaccard overlap 51% of words appear in one view only
Cross-modal grounding Independently produced captions and poses agree on when the camera turns 0.709 AUC yaw 1.3° → 27.9° across labels
Within-clip control The ordering survives when each clip is its own control 71.8% of clips median gap 12.7° · 74,437 clips
Frame quality Long footage is not bought with worse images 5.23 aesthetic MiraData 5.45 · RE10K 5.20
Frame-to-frame motion Sustained translation rather than slow orbits around a static scene 35.1 flow 1.4–1.8× the prior corpora

Build with Sekai2

The world is larger
than a short clip.

The technical report, dataset access, annotation code and benchmark splits will appear here. Every clip ships video, a ViPE camera trajectory and a hierarchical caption under one schema.