The Internet's Hidden 3D Model of the World

Sumber Unduh video ini
en-origen
May 14, 2026 Jul 14, 2026
Video preview
Bagikan:

Recent advances in computer vision now allow us to fuse billions of random internet photos into coherent 3D models, turning the long tail of poorly photographed places into navigable digital twins.

From Structure from Motion to Gaussian Splatting ⏱ 0:00

  • 2009: University of Washington downloads thousands of tourist photos of Rome from Flickr, uses structure from motion to reverse-engineer camera positions and stitch photos into a 3D model ("Building Rome in a Day").
  • Techniques from that paper (skeletal sets, bundle adjustment at scale, posing photos against global 3D model) became the toolkit still used by Street View.
  • By 2015, a UNC team reconstructed the entire planet from Flickr in 6 days.
  • Problem: Photos online are unevenly distributed — only famous landmarks (the "head") have many photos; the rest ("long tail") have few, leading to hollow reconstructions.
  • 2021: Google's NeRF in the Wild adapts neural radiance fields to internet-scale photos, enabling disentanglement of lighting and scene (e.g., changing time of day on Brandenburg Gate).
  • 2023: 3D Gaussian Splatting replaces implicit neural networks with explicit fuzzy ellipsoidal splats, enabling real-time rendering at 100+ FPS in a browser.
  • 2024: Wild Gaussians applies the "in the wild" trick to the new substrate, allowing interactive sunrise-to-sunset sliders in a browser.
  • Solving the Long Tail Problem ⏱ 9:00

  • The Doppelgangers problem: bilateral symmetry (e.g., identical front and back of a building) causes reconstructions to fold. Noah Snavely's team trains a transformer (doppelgangers++) to detect these.
  • Feed-forward models like VGGT (CVPR 2025 Best Paper) and Pi-Cube predict cameras and geometry from a pile of photos in seconds, bypassing structure from motion.
  • Pi-Cube fixes VGGT's weakness: VGGT secretly picks one reference photo; Pi-Cube removes the anchor entirely.
  • April 2026: MegaDepth X by Cornell (Noah Snavely, PhD student Wan Li) tackles the chicken-and-egg problem of sparse internet photos.
  • - Trick: take well-photographed landmarks with ground truth 3D reconstructions, throw away most photos to simulate the long tail, and train on that "hard problem with a stolen answer key."

    - Fine-tuning VGGT and Pi-Cube on MegaDepth X improves rotation accuracy on hardest sparse scenes from 75% (off-the-shelf Pi-Cube) to 86%.

    Fusing Multi-Source Data and Military Applications ⏱ 11:38

  • Diffusion models fill gaps where data is sparse. Example: Sky Fall GS uses 3D Gaussian Splatting plus diffusion models to fill satellite gaps (solving Google Earth's biggest problem from above).
  • April 2025: SRI International's Diffusion Guided Gaussian Splatting fuses ground-level photos, drone shots, and satellite data into one 3D model, letting diffusion models fill any source's gaps.
  • SRI is contractor for IARPA's WRIVA project (Walkthrough Rendering from Images of Varying Altitudes), a 42-month effort since 2023 to build photorealistic 3D walkthroughs of places agents cannot physically go (the intelligence community's long-tail problem).
  • MegaDepth X funded by Korea's National AI Research Lab Project — same race, different country.
  • Patterns: same tech appears in Cornell research, Netflix VFX tools, and intelligence programs within months.
  • The Road to 4D and the Big Picture ⏱ 14:13

  • World is 4D (people move, time passes). Papers like Mosca and Shape-a-Motion can pull 4D (geometry + motion) from a single phone video clip.
  • Not yet at fusing every concertgoer's iPhone into free-viewpoint playback, but direction is clear.
  • Every iPhone, dashcam, and online photo can now be used to extract 3D structure. The sensorium has come to life — a real God's eye view built from vacation photos.
  • Until last month, it was hard to pull together; now all viewpoints can be molded into a 3D view.
  • Key Takeaways

  • 2009's 'Building Rome in a Day' pioneered using tourist photos for 3D reconstruction; techniques still underpin Street View.
  • The long-tail problem (places with few photos) was the main barrier; MegaDepth X (April 2026) overcomes it by simulating sparse data.
  • Feed-forward models (VGGT, Pi-Cube) now predict cameras and geometry in seconds, replacing slow structure from motion.
  • Fusing ground, drone, and satellite data with diffusion models fills gaps; IARPA's WRIVA project uses this for military reconnaissance.
  • Current work is extending from static 3D to 4D (motion from handheld video), aiming to reconstruct dynamic scenes.
  • Conclusion

    We now have the ability to turn any collection of internet photos into a coherent 3D model, unlocking the long tail of the world. This technology has both commercial and military applications, and the race is accelerating globally.

    Tanya AI tentang video ini

    Sorotan Visualbeta