Best LLM Papers

Best LLM Papers · 2026-W40

Ranking window: 2026-07-06 – 2026-10-03 Frozen on 2026-10-04

4

LingBot-VLA 2.0: a vision-language-action model pretrained on 60,000 hours spanning 20 robot embodiments

This technical report upgrades LingBot-VLA with a revamped data pipeline of about 60,000 pretraining hours: 50K robot-trajectory hours across 20 embodiments plus 10K egocentric human-video hours. It extends control beyond dual arms to heads, waists, mobile bases and dexterous hands, aiming squarely at the gap between lab robotics and real deployment.

  • 28 stars/7d
  • 15 citations
  • 21 upvotes
  • #3 robotwin-2-0-easy-50-tasks
5

Xiaomi Robotics 1: a vision-language-action model pretrained on over 100K hours of real-world trajectories

Xiaomi's robotics team presents a vision-language-action foundation model for mobile manipulation, pretrained on more than 100K hours of real-world UMI trajectories with auto-labeled scene-transition language, then post-trained to match embodiments and human imperative prompts. The paper claims clean scaling with more data and parameters, out-of-the-box performance in unseen environments and efficient fine-tuning for dexterous tasks.

  • 8 stars/7d
  • 15 citations
  • 75 upvotes
  • #3 robocasa
6

YuE2: one Mixture-of-Transformers that writes a readable score first, then renders full-song audio from it

Symbolic music models plan composition explicitly but never sound like a finished recording, while audio models produce full songs that keep the composition hidden. YuE2 combines both in a single AR-NAR Mixture-of-Transformers that plans symbolically, expands the plan into semantic music tokens, and renders the full song, using two new representation models (MERT2, SheetSage2) to learn from recordings that lack aligned scores. The authors report a SongBench Global Avg of 6.73 on WildSongBench (6.96 at best-of-8), and in expert listening the best-of-8 samples were preferred over Suno v4.5 and roughly tied with Suno v5; the readable score also makes the output editable by external language-model agents.

  • 431 stars/7d
  • 0 citations
  • 216 upvotes
7

RoboDojo: one benchmark that puts generalist robot policies through 42 simulation and 18 real-world tasks

RoboDojo is a unified sim-and-real benchmark for generalist robot manipulation policies, with 42 simulation tasks and 18 real-world tasks probing generalization, memory, precision and long-horizon control. It ships parallel Isaac Sim evaluation plus a cloud-accessible real-eval system and a leaderboard the authors populated with 30 policies.

  • 31 stars/7d
  • 12 citations
  • 17 upvotes
8

WROP: a cognitive-science exam that tests whether video world models understand object permanence — and a 16B model trained to pass it

When an object leaves the frame, does a video generation model remember it exists? The authors built WROP, 150 cognitive-science tasks across six categories rendered in Blender with randomized lighting, speed, and camera angle — a 1.5M-sample training corpus plus a 300-question exam used to test 14 video models. They then trained PWM-WROP, a 16B world model on their native PyTorch stack for AWS Trainium2, releasing data, exam, scores, and weights. In a blind pairwise Elo study, it ranked first among continuation models and effectively tied the best reference-to-video models.

  • 303 stars/7d
  • 0 citations
  • 193 upvotes
9

Frontis-MA1: an open 35B meta-evolution agent for recursive self-improvement in ML engineering

Frontis-MA1 ships with OpenMLE, an open full-stack system (task gyms, RL operator learning, evolutionary search) for studying recursive self-improvement. Its 35B meta-evolution agent, post-trained around Draft, Improve, Debug and Crossover program-evolution operators, claims MLE-Bench Lite results the authors say exceed GPT-5.5 plus Codex, with components transferring to held-out NatureBench Lite.

  • 23 stars/7d
  • 3 citations
  • 186 upvotes
13

Edge0 streams 35B MoE experts from SSD, predicting each layer's routing one token ahead

Edge0 is an open-source inference engine that serves a 35B mixture-of-experts model from a single 24GB GPU by loading expert weights from SSD on demand; a trained per-layer prerouter predicts the next layer's expert choices one token ahead so SSD reads overlap with compute, and a recovery LoRA pays back the quality lost to int4 quantization. The authors report 20 tok/s at about 3GiB of peak active memory, within a few points of the fp16 teacher on five public benchmarks.

  • 664 stars/7d
  • 0 citations
  • 8 upvotes
14

Spark-to-Paper: from research idea to submission-ready paper, as thirteen composable skills

Spark-to-Paper implements end-to-end paper generation as thirteen composable skills inside an existing coding assistant, with no separate agent platform: it retrieves literature, designs and runs experiments, revises claims to match evidence and produces figures, keeping model judgment separate from mechanically checkable operations.

  • 69 stars/7d
  • 0 citations
  • 291 upvotes
15

OneStreamer: a streaming video LLM that keeps a written memory of what it saw and speaks up when the evidence is there

Streaming video assistants must watch a live feed in real time and still remember what happened minutes ago. OneStreamer (Nanjing University) trains a 4B model to do both with one shared generation process: while watching, it proactively writes time-stamped captions and summaries into a hierarchical text memory, so follow-up questions are answered from those written records instead of re-encoding old frames. It also releases OneStreamer-1M, a one-million-record dataset for streaming video interaction, and reports best results among compared methods on all eight evaluated streaming understanding benchmarks.

  • 98 stars/7d
  • 0 citations
  • 154 upvotes

This week · Best LLM Papers