// measuring visual intelligence_

Physics-IQ Verified

GitHub repoCode & benchmarkFull paperMethods & resultsDataset auditFixes & visual evidence
Benchmark track

I2V Verified Ranking

Model results

1
48.2%±1.4
i2vYes2026-09-25
2
43.3%±1.5
i2vYes2026-09-25
3
42.7%±0.8
i2vYes2026-09-04
4
42.4%±0.8
i2vYes2026-09-28
5
39.8%±0.3
i2vYes2026-08-24
6
37.3%±0.9
i2vYes2026-09-04
7
36.2%±0.7
i2vYes2026-08-27
8
35.3%±0.4
i2vYes2026-09-28
9
34.8%±0.6
i2vYes2026-06-17
10
33.7%±1.4
i2vYes2026-06-19
11
33.4%±0.8
i2vNo2026-06-17
12
33.4%±0.2
i2vYes2026-09-28
13
32.7%±1.1
i2vYes2026-08-31
14
32.2%±0.6
i2vNo2026-06-17
15
31.8%±0.3
i2vYes2026-09-28
16
31.8%±1.5
i2vNo2026-08-18
17
30.8%±0.9
i2vYes2026-08-07
18
30.0%±0.5
i2vYes2026-09-28
19
27.7%±0.9
i2vNo2026-08-18
20
26.5%±0.8
i2vYes2026-06-17
21
25.3%±1.8
i2vNo2026-06-17

Metric Breakdown

Submetric leaders

Spatial

  1. Physis-Lang (Cosmos3 Super)59.9±1.4
  2. MiniMax H358.9±0.8
  3. Cosmos3 Super57.4±1.5
  4. Seedance 2.557.2±0.7
  5. Physis-Lang (Cosmos3 Nano)56.2±1.8

Spatiotemporal

  1. Physis-Lang (Cosmos3 Super)41.6±3.1
  2. CogVideoX-5B35.5±4.9
  3. Cosmos3 Super32.5±1.0
  4. Physis-Lang (Cosmos3 Nano)31.0±2.2
  5. Kandinsky-WM 1.030.0±2.7

Weighted Spatial

  1. Physis-Lang (Cosmos3 Super)48.5±1.1
  2. Physis-Lang (Cosmos3 Nano)43.4±2.2
  3. Cosmos3 Super43.0±1.5
  4. Seedance 2.542.7±0.9
  5. MiniMax H340.5±0.8

MSE

  1. Physis-Lang (Cosmos3 Super)43.0±1.3
  2. Physis-Lang (Cosmos3 Nano)42.6±0.8
  3. Seedance 2.541.3±1.0
  4. Cosmos3 Super37.8±0.4
  5. MiniMax H337.0±0.7

Cost Frontier

Score vs Cost ($)

* Price via leading API providers or estimated via GPU market rate, May 2026. Generation cost is normalized to 24 FPS and 1280-wide output. Separate LLM prompt overhead is added after video normalization where used. n.d. denotes values not publicly disclosed by the model provider. GPU implementations were done to the best of our knowledge and as close as possible to the recommended setup.