Skip to content

Latest commit

 

History

251 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Video Depth Estimation Rankings
and Stereo Video Conversion Rankings

Table of Contents

Introduction

Stereo Video Conversion Rankings

Video Depth Estimation Rankings

Single Image Depth Estimation Rankings

Appendices


Purpose of this repository

"The 3D shows you a window into reality; the higher frame rate takes the glass out of the window"
— James Cameron

Researchers, if you have found your way here, please consider developing a new Stereo Video Conversion model based on MiniMax H3.

Back to Table of Contents

Awesome Stereo Video Conversion

The following list includes all Stereo Video Conversion methods from the last 16 months, from 1 April 2025 to 1 August 2026. This list was created because there is a significant problem with public access to the latest Stereo Video Conversion models, which makes it difficult for researchers to compare their work with the current state of the art and to use the same test set. Consequently, this also makes it difficult to present a single ranking that showcases all of the best models.

Method Backbone Submitted on
(arXiv)
     Venue      Official
  repository  
Code
(website)
GRT ProPainter 6 Jul 2026 ICML - -
αDepth ? 29 May 2026 arXiv - -
StereoCrafter2 Wan2.1-VACE-14B - - GitHub Stars -
DreamStereo Wan2.1-1.3B 14 Apr 2026 CVPR - GitHub Stars
HairGuard Wan2.1-VACE-1.3B 6 Jan 2026 arXiv - -
StereoPilot Wan2.1-T2V-1.3B 18 Dec 2025 arXiv GitHub Stars GitHub Stars
Elastic3D SVD 16 Dec 2025 CVPR - GitHub Stars
StereoWorld Wan2.1-T2V-1.3B 10 Dec 2025 CVPR GitHub Stars GitHub Stars
Restereo StereoCrafter (based on SVD) 6 Jun 2025 CVPRW - -
M2SVid SVD 22 May 2025 3DV GitHub Stars GitHub Stars
Eye2Eye Lumiere 30 Apr 2025 arXiv - GitHub Stars

Back to Table of Contents

Awesome Synthetic RGB-D Image Datasets for Training HD Video Depth Estimation Models

Although video depth estimation models should be trained mainly on synthetic RGB-D video datasets I decided to add two synthetic RGB-D image datasets because of their unique features.

Dataset      Venue      Resolution Unique features
1 SynthHuman
📌 Human faces 😍
ICCV 384×512 The dataset contains 98040 samples feature the face, 99976 sample feature the full body and 99992 samples feature the upper body. DAViD trained on this dataset alone achieved better depth estimation results than Depth Anything V2 Large, Depth Pro and even Sapiens-2B on the Goliath-Face test set. See the results in Table 1.
2 MegaSynth CVPR 512×512 Huge size: 700K scenes and the incredible improvement in depth estimation results of the fine-tuned Depth Anything V2 ViT-B model on MegaSynth and evaluated on Hypersim. See the results in Table 6.

Back to Table of Contents

Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

The following list contains only synthetic RGB-D datasets in which at least some of the images can be composited into a video sequence of at least 32 frames. The minimum number of frames was chosen on the basis of the ablation studies shown in Table 5 by the Video Depth Anything researchers.

Most datasets contain ready-to-use video sequences of appropriately numbered images in individual folders, but in the case of the PLT-D3 dataset, images from at least two folders have to be combined to make a longer video sequence and in the case of the ClaraVid dataset, images have to be arranged in the correct order to make a 32-frame video sequence, for example in the order given in Appendix 4.

Researchers, if you are going to use the following list to select datasets to train your models check their quality very carefully and choose the best ones. I have only visually checked a few of them and have marked on the list 2 datasets to check particularly carefully and 2 datasets that in my opinion are not suitable for training video depth estimation models. I have given the reasons for such markings in the same Appendix 4.

In selecting the best datasets, comparisons of their quality can be very helpful, such as in Table 9, Table 6, another Table 6 for depth estimation models and TABLE V plus TABLE IV for stereo matching models, although a similar technique can also be used for depth estimation models.

Dataset        Venue        Resolution V
G
D
A
3
T
M
o
3
M
o
2
D
P
U
D
2
G
D
V
D
A
D
V
D
D
C
R
D
1 SynthVerse
📌 Human face 4K 😍

in the 164-frame battery-grabbing scene from "Charge"
SIGGRAPH 4096×1716 - - - - - - - - - - -
2 BEDLAM2.0
📌 Human poses 😍
NeurIPS 1280×720 - - - - - - - - - - -
3 C3I-SynFace
📌 Human faces 😍
DIB 640×480 - - - - - - - - - - -
4 ClaraVid ICCV 4032x3024 - - - - - - - - - - -
5 StereoGenBench arXiv 1280×1280 - - - - - - - - - - -
6 Spring CVPR 1920×1080 T T E E E - - - - - -
7 HorizonGS CVPR 1920×1080 - - - - - - - - - - -
8 PLT-D3 HD 1920×1080 - - - - - - - - - - -
9 MVS-Synth CVPR 1920×1080 T T T T T - T - - - -
10 SYNTHIA-SF BMVC 1920×1080 - - - - - - - - - - -
11 SynDrone
Check before use!
ICCVW 1920×1080 - - - - - - - - - - -
12 Mid-Air CVPRW 1024×1024 - - T T - - - - - - -
13 MatrixCity ICCV 1000×1000 T T T T - T - - - T -
14 StereoCarla arXiv 1600×900 - - - - - - - - - - -
15 LightwheelOcc - 1600×900 T - - - - - - - - - -
16 SAIL-VOS 3D CVPR 1280×800 - - - - T - - - - - -
17 SHIFT CVPR 1280×800 - - - - - - - - - - -
18 SYNTHIA-Seqs
🚫 Do not use! 🚫
CVPR 1280×760 T - T T - - - - - - -
19 BEDLAM CVPR 1280×720 T - - - T T - - - - -
20 Dynamic Replica CVPR 1280×720 T - - - T T T - - T -
21 OmniWorld-Game ICLR 1280×720 T - T - - - - - - - -
22 WorldRover arXiv 1280×720 - - - - - - - - - - -
23 Syn4D ECCV 1280×720 - - - - - - - - - - -
24 InFlux++ Synth ECCV 1280×720 - - - - - - - - - - -
25 Infinigen SV TPAMI 1280×720 - - - - - - - - - - -
26 Infinigen CVPR 1280×720 - - - - - - - - - - -
27 DigiDogs
🚫 Do not use! 🚫
WACVW 1280×720 - - - - - - - - - - -
28 Aria Synthetic Environments
Check before use!
- 704×704 T - - - - - - - - - -
29 TartanGround IROS 640×640 T - - - - - - - - - -
30 TartanAir V2 - 640×640 - - - - - - - - - - -
31 BlinkVision ECCV 960×540 - - - - - - - - - - -
32 PointOdyssey ICCV 960×540 T T - - - T T T - - E
33 DyDToF CVPR 960×540 - - - - - - - - - - E
34 IRS ICME 960×540 - T T T T - T T - - -
35 Scene Flow CVPR 960×540 - - - - - - - - - - -
36 THUD++ arXiv 730×530 - - - - - - - - - - -
37 TAU Agent TCI 1024×512 - T - - - - - - - - -
38 TransPhy3D ICRA 512×512 T - - - - - - - - - -
39 3D Ken Burns TOG 512×512 - T T T T - - - - - -
40 SynPhoRest - 848×480 - - - - - - - - - - -
41 TartanAir IROS 640×480 T T T T T T T T T - T
42 ParallelDomain-4D ECCV 640×480 - - - - - - - - - - -
43 EDEN WACV 640×480 - T T T T T - - - - -
44 GTA-SfM RAL 640×480 T T T T - - - - - - -
45 InteriorNet BMVC 640×480 - - - - - - - - - - -
46 SYNTHIA-AL ICCVW 640×480 - - - - - - - - - - -
47 MPI Sintel ECCV 1024×436 E E E E E E E E E E -
48 CarlaOcc CVPR 1408×376 T - - - - - - - - - -
49 Virtual KITTI 2 arXiv 1242×375 - T - - T - T T - - -
50 Virtual KITTI CVPR 1242×375 - - - - - - - - T - -
51 TartanAir Shibuya ICRA 640×360 - - - - - - - - - - -
Total: T (training) 15 11 10 9 9 6 6 4 2 2 1
Total: E (testing) 1 1 2 2 2 1 1 1 1 1 2

Back to Table of Contents

Stereo4D (400 video clips with 16 frames each at 5 fps): LPIPS<=0.242

RK Model
Links:
         Venue   Repository    
   LPIPS ↓   
{Input fr.}
3DV
Table 1
M2SVid
1 M2SVid
3DV GitHub Stars
0.180 {MF}
2 SVG
ICLR GitHub Stars
0.217 {MF}
3 StereoCrafter
arXiv GitHub Stars
0.242 {MF}

Back to Table of Contents

StereoWorld-11M (1000 video clips with 81 frames each at 12 fps): LPIPS<=0.1869

RK Model
Links:
         Venue   Repository    
   LPIPS ↓   
{Input fr.}
CVPR
Table 2
StereoWorld
1 StereoWorld
CVPR GitHub Stars
0.0952 {MF}
2 StereoCrafter
arXiv GitHub Stars
0.1869 {MF}

Back to Table of Contents

170-frame ScanNet: TAE

📝 Note: This ranking is based on the evaluation protocol proposed by Video Depth Anything developers.

RK Model
Links:
         Venue   Repository    
  TAE ↓  
{Input fr.}
ECCV
Table 3
FFN
  TAE ↓  
{Input fr.}
ICML
Table 2&
GitHub Stars
GD
  TAE ↓  
{Input fr.}
CVPR
Table 1
VDA
1 FFN
ECCV GitHub Stars
0.380 {MF} - -
2 GemDepth-VDA
ICML GitHub Stars
- 0.47 {MF} -
3 VDA-L
CVPR GitHub Stars
-
VDA-L-Syn:
0.570 {MF}
0.57 {MF}
VDA-L-Syn:
-
0.570 {MF}
VDA-L-Syn:
0.570 {MF}
4 DVD v1.1
ICMLW GitHub Stars
- 0.61 {MF} -
5 DepthCrafter
CVPR GitHub Stars
0.639 {MF} - 0.639 {MF}
6 RollingDepth
CVPR GitHub Stars
- 0.65 {MF} -
7 Depth Any Video
ICLR GitHub Stars
0.967 {MF} - 0.967 {MF}
8 ChronoDepth
CVPR GitHub Stars
1.022 {MF} - 1.022 {MF}
9 Depth Anything V2 Large
NeurIPS GitHub Stars
1.140 {1} 1.14 {1} 1.140 {1}
10 NVDS
ICCV GitHub Stars
2.176 {4} - 2.176 {4}

Back to Table of Contents

500-frame Bonn RGB-D Dynamic: δ1

📝 Note 1: This ranking is based on the evaluation protocol proposed by Video Depth Anything developers.
📝 Note 2: A high rank of the Pixel-Perfect Video Depth model in this ranking does not guarantee that this model is suitable for practical applications, see gangweix/pixel-perfect-depth#27. We are waiting for this model to be fixed and evaluated using temporal stability metrics.

RK Model
Links:
         Venue   Repository    
     δ1 ↑     
{Input fr.}
ECCV
Table 3
FFN
     δ1 ↑     
{Input fr.}
arXiv
TABLE II
PPVD
     δ1 ↑     
{Input fr.}
ICML
Table 1&
GitHub Stars
GD
     δ1 ↑     
{Input fr.}
CVPR
Table 1
VDA
1 FFN
ECCV GitHub Stars
0.982 {MF} - - -
2 Pixel-Perfect Video Depth
arXiv GitHub Stars
- 0.979 {MF} - -
3 GemDepth-VDA
ICML GitHub Stars
- - 0.978 {MF} -
4 VDA-L
CVPR GitHub Stars
-
VDA-L-Syn:
0.961 {MF}
0.959 {MF}
VDA-L-Syn:
-
0.959 {MF}
VDA-L-Syn:
-
0.959 {MF}
VDA-L-Syn:
0.961 {MF}
5 DVD v1.1
ICMLW GitHub Stars
- - 0.948 {MF} -
6 RollingDepth
CVPR GitHub Stars
- 0.931 {MF} 0.931 {MF} -
7 Depth Anything V2 Large
NeurIPS GitHub Stars
0.864 {1} 0.864 {1} 0.864 {1} 0.864 {1}
8 DepthCrafter
CVPR GitHub Stars
0.803 {MF} 0.803 {MF} 0.803 {MF} 0.803 {MF}
9 NVDS
ICCV GitHub Stars
0.674 {4} 0.674 {4} 0.674 {4} 0.674 {4}
10 ChronoDepth
CVPR GitHub Stars
0.665 {MF} 0.665 {MF} 0.665 {MF} 0.665 {MF}

Back to Table of Contents

Synth4K: δ1

RK Model
Links:
         Venue   Repository    
     δ1 ↑     
{Input fr.}
arXiv
Table C.1
MoGe-3
1 MoGe-3 ViT-G Step 3
arXiv GitHub Stars
0.949 {1}
2 MoGe-2
NeurIPS GitHub Stars
0.913 {1}
3 Depth Anything 3
ICLR GitHub Stars
0.883 {1}
4 InfiniDepth
CVPR GitHub Stars
0.878 {1}
5 UniDepthV2
arXiv GitHub Stars
0.869 {1}
6 Pixel-Perfect Depth
NeurIPS GitHub Stars
0.868 {1}
7 Depth Pro
ICLR GitHub Stars
0.866 {1}
8 UniK3D
CVPR GitHub Stars
0.862 {1}

Back to Table of Contents

Appendix 4: Notes on the table: Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

📝 Note 1: Example of arranging images in the correct order to make a 32-frame video sequence for the ClaraVid dataset:

<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00360.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00320.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00280.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00240.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00200.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00160.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00120.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00080.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00040.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00000.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00001.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00002.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00003.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00004.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00005.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00006.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00007.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00008.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00009.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00010.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00011.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00012.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00013.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00014.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00015.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00016.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00017.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00018.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00019.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00059.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00099.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00139.jpg

📝 Note 2: Do not use the SYNTHIA-Seqs dataset for training HD video depth estimation models! The depth maps in this dataset do not match the corresponding RGB images. This is particularly evident in the example of tree leaves:
<your-data-path>/SYNTHIA-SEQS-01-SPRING/Depth/Stereo_Left/Omni_F/000071.png
<your-data-path>/SYNTHIA-SEQS-01-SPRING/RGB/Stereo_Left/Omni_F/000071.png.
📝 Note 3: Do not use the DigiDogs dataset for training HD video depth estimation models! The depth maps in this dataset do not match the corresponding RGB images. See the objects behind the campfire, the shifting position of the vegetation on the left and the clear banding on the depth map:
<your-data-path>/DigiDogs2024_full/09_22_2022/00054/images/img_00012.tiff.
📝 Note 4: Check before use the SynDrone dataset for training HD video depth estimation models! The depth maps in this dataset have large white areas of unknown depth, which should not happen with a synthetic dataset. Example depth map:
<your-data-path>/Town01_Opt_120_depth/Town01_Opt_120/ClearNoon/height20m/depth/00031.png.
📝 Note 5: Check before use the Aria Synthetic Environments dataset for training HD video depth estimation models! The depth maps in this dataset have large white areas of unknown depth, which should not happen with a synthetic dataset. Example depth map:
<your-data-path>/75/depth/depth0000109.png.

Back to Table of Contents

Appendix 5: List of all research papers from the above rankings

Method Abbr. Paper      Venue     
(Alt link)
Official
  repository  
ChronoDepth - Learning Temporally Consistent Video Depth from Video Diffusion Priors CVPR GitHub Stars
Depth Any Video DAV Depth Any Video with Scalable Synthetic Data ICLR GitHub Stars
Depth Anything 3 DA3 Depth Anything 3: Recovering the Visual Space from Any Views ICLR GitHub Stars
Depth Anything V2 DA V2 Depth Anything V2 NeurIPS GitHub Stars
Depth Pro DP Depth Pro: Sharp Monocular Metric Depth in Less Than a Second ICLR GitHub Stars
DepthCrafter DC DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos CVPR GitHub Stars
DVD - DVD: Deterministic Video Depth Estimation with Generative Priors ICMLW GitHub Stars
FFN - Forget, Anticipate and Adapt: Test Time Training for Long Videos ECCV GitHub Stars
GemDepth GD GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth ICML GitHub Stars
InfiniDepth - InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields CVPR GitHub Stars
M2SVid - M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion 3DV GitHub Stars
MoGe-2 Mo2 MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details NeurIPS GitHub Stars
MoGe-3 Mo3 MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement arXiv GitHub Stars
NVDS - Neural Video Depth Stabilizer ICCV GitHub Stars
Pixel-Perfect Depth PPD Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers NeurIPS GitHub Stars
Pixel-Perfect Video Depth PPVD Pixel-Perfect Visual Geometry Estimation arXiv GitHub Stars
RollingDepth RD Video Depth without Video Models CVPR GitHub Stars
StereoCrafter - StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos arXiv GitHub Stars
StereoWorld - StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation CVPR GitHub Stars
SVG - SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix ICLR GitHub Stars
UniDepthV2 UD2 UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler arXiv GitHub Stars
UniK3D - UniK3D: Universal Camera Monocular 3D Estimation CVPR GitHub Stars
Video Depth Anything VDA Video Depth Anything: Consistent Depth Estimation for Super-Long Videos CVPR GitHub Stars

Back to Table of Contents

Appendix 6: List of all research papers from the column headers of the table: Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

Method Abbr. Paper      Venue     
(Alt link)
Official
  repository  
Depth Anything 3 DA3 Depth Anything 3: Recovering the Visual Space from Any Views ICLR GitHub Stars
Depth Pro DP Depth Pro: Sharp Monocular Metric Depth in Less Than a Second ICLR GitHub Stars
DepthCrafter DC DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos CVPR GitHub Stars
DVD - DVD: Deterministic Video Depth Estimation with Generative Priors ICMLW GitHub Stars
GemDepth GD GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth ICML GitHub Stars
MoGe-2 Mo2 MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details NeurIPS GitHub Stars
MoGe-3 Mo3 MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement arXiv GitHub Stars
RollingDepth RD Video Depth without Video Models CVPR GitHub Stars
UniDepthV2 UD2 UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler arXiv GitHub Stars
Video Depth Anything VDA Video Depth Anything: Consistent Depth Estimation for Super-Long Videos CVPR GitHub Stars
ViGeo VG Towards Consistent Video Geometry Estimation arXiv GitHub Stars

Back to Table of Contents

List of research papers to be added to the rankings

Method Abbr. Paper      Venue     
(Alt link)
Official
  repository  
αDepth - αDepth: Learning Single-Pass Soft Boundary Decomposition for Stereo Conversion arXiv -
DreamStereo - DreamStereo: Towards Real-Time Stereo Inpainting for HD Videos CVPR -
GRT - Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation ICML -
HairGuard - Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel Views arXiv -
StereoPilot - StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors arXiv GitHub Stars
Elastic3D - Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding CVPR -
Restereo - Restereo: Unifying diffusion stereo video generation and restoration CVPRW -
Eye2Eye - Eye2Eye: A Simple Approach for Monocular-to-Stereo Video Synthesis arXiv -

Back to Table of Contents

About

Researchers, we look forward to models based on MiniMax H3. ChronoDepth Depth Any Video Depth Anything Depth Pro DepthCrafter DreamStereo DVD Elastic3D Eye2Eye FFN GemDepth GRT HairGuard InfiniDepth MoGe M2SVid NVDS Pixel-Perfect Video Depth Restereo StereoCrafter StereoPilot StereoWorld SVG UniDepth UniK3D Video Depth Anything ViGeo αDepth

Topics

Resources

Stars

259 stars

Watchers

14 watching

Forks

Packages

Contributors