AI research papers explained visually

Understand the method, evidence, and limits—not just the abstract

Publicly shareable paper PDFs only (25 MB max)
View an English example

Research papers explained in plain English

Robotics / NVIDIA / English

Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control

Jun Yamada, Adithyavairavan Murali, Ajay Mandlekar, Clemens Eppner, Ingmar Posner, and Balakumar Sundaralingam · ICRA 2026 · arXiv:2509.06201

arXiv 2509.062017/26/2026
An open-loop robot follows an outdated grasp target after the object shifts.
An open-loop robot follows an outdated grasp target after the object shifts.
Robotics / NVIDIA / English

Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

Wenli Xiao et al. · ICLR 2026 conference paper · Primary paper

arXiv 2511.000917/26/2026
Trending nowA frozen generalist enters failure regions that polished demonstrations rarely cover
A frozen generalist enters failure regions that polished demonstrations rarely cover
Robotics / NVIDIA / English

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Shenyuan Gao et al. · ICML 2026 Spotlight · arXiv:2602.06949

arXiv 2602.069497/26/2026
The human-video mixture is far larger and more varied than robot datasets used by earlier world models.
The human-video mixture is far larger and more varied than robot datasets used by earlier world models.
Robotics / NVIDIA / English

FLARE: Robot Learning with Implicit World Modeling

Ruijie Zheng et al. · Conference on Robot Learning (CoRL) 2025 · Paper · Proceedings

arXiv 2505.156597/26/2026
A comparison between predicting detailed future pixels and compact future features
A comparison between predicting detailed future pixels and compact future features
Robotics / NVIDIA / English

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

Joel Jang et al. · Conference on Robot Learning (CoRL) 2025 · Paper · Proceedings

arXiv 2505.127057/26/2026
A narrow real-data seed teaches one robot body
A narrow real-data seed teaches one robot body
Robotics / NVIDIA / English

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

NVIDIA technical report on arXiv (arXiv:2503.14734, version 2, 27 March 2025). This is not presented as a peer-reviewed conference paper.

arXiv 2503.147347/26/2026
A pyramid connects broad human experience to scarce robot demonstrations
A pyramid connects broad human experience to scarce robot demonstrations
Robotics / Humanoid / English

ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills

Tairan He et al. · arXiv:2502.01143v3 · 2025 · Primary source

arXiv 2502.011437/26/2026
A human motion passes through reconstruction, physics screening, and body-shape retargeting before it becomes a robot training target.
A human motion passes through reconstruction, physics screening, and body-shape retargeting before it becomes a robot training target.
Robotics / Humanoid / English

HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots

Tairan He et al. · arXiv:2410.21229v2 · ICRA 2025 · Primary paper

arXiv 2410.212297/26/2026
Human motion is matched to robot body points, filtered for feasibility, and used for oracle practice.
Human motion is matched to robot body points, filtered for feasibility, and used for oracle practice.
Robotics / Humanoid / English

OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning

Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris Kitani, Changliu Liu, and Guanya Shi · arXiv:2406.08858 · submitted 13 June 2024 · paper

arXiv 2406.088587/26/2026
Human motion is retargeted, filtered, and augmented with stable poses
Human motion is retargeted, filtered, and augmented with stable poses
AI / Technology / English

LingBot-World: an open-source video world simulator with long memory and real-time control

This paper argues that video generators can be pushed beyond short clips into an interactive world simulator: one that keeps track of what happened earlier, reacts to user actions, and still moves fast enough to feel live.

arXiv 2601.205407/23/2026
A white-canvas illustration showing two side-by-side video sequences of the same scene: one sequence stays coherent as a ball rolls, a person walks, and a door remains where it was; the other sequence looks visually plausible frame by frame but objects change position inconsistently, the ball disappears, and the door location drifts. Blue arrows indicate cause and effect across time. Short labels only: "Dreamer", "Simulator", "Memory", "Cause", "Effect".
A believable clip versus a living world
AI / Technology / English

ImageNet Classification with Deep Convolutional Neural Networks

Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · NeurIPS 2012

proceedings.neurips.cc7/19/2026
A left-to-right teaching diagram of the network, from an image through five convolutional layers and three fully connected layers to 1,000 classes.
A left-to-right teaching diagram of the network, from an image through five convolutional layers and three fully connected layers to 1,000 classes.
AI / Technology / English

Segment Anything

Paper: Alexander Kirillov et al. · 2023 · arXiv:2304.02643 · Full paper

arXiv 2304.026437/19/2026
A single point can reasonably refer to a badge, jacket, or whole person.
A single point can reasonably refer to a badge, jacket, or whole person.
AI / Technology / English

Visual Instruction Tuning

Authors: Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee Published: NeurIPS 2023 (Oral); arXiv v2, 11 December 2023 Primary source: Paper abstract · Full paper (PDF) Reading note: Results and limitations

arXiv 2304.084857/19/2026
Flow from image annotations through text-only GPT-4 to three kinds of instruction data
Flow from image annotations through text-only GPT-4 to three kinds of instruction data
AI / Technology / English

ReAct: Synergizing Reasoning and Acting in Language Models

Authors: Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao Published: ICLR 2023 · arXiv v3, 10 March 2023 Primary source: arXiv:2210.03629 · PDF Reading note: Results belo

arXiv 2210.036297/19/2026
The ReAct cycle: thought guides action, and observations revise the next thought.
The ReAct cycle: thought guides action, and observations revise the next thought.
AI / Technology / English

Training Compute-Optimal Large Language Models

Jordan Hoffmann and colleagues · arXiv:2203.15556 · submitted 29 March 2022

arXiv 2203.155567/19/2026
A fixed compute budget split between parameters and training tokens, with balanced scaling leading toward lower loss.
A fixed compute budget split between parameters and training tokens, with balanced scaling leading toward lower loss.
AI / Technology / English

Training language models to follow instructions with human feedback

Long Ouyang et al. · OpenAI · submitted 4 March 2022 · arXiv:2203.02155 · PDF

arXiv 2203.021557/19/2026
Two different goals: predicting internet text and serving a user's intent
Two different goals: predicting internet text and serving a user's intent
AI / Technology / English

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Paper: Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · arXiv:2201.11903 · submitted 28 January 2022, revised 10 January 2023 · primary source · [full text](https://ar5iv.labs.arxiv.org/ht

arXiv 2201.119037/19/2026
Two prompting lanes: standard examples lead straight to an answer, while chain-of-thought examples include intermediate steps.
Two prompting lanes: standard examples lead straight to an answer, while chain-of-thought examples include intermediate steps.
AI / Technology / English

High-Resolution Image Synthesis with Latent Diffusion Models

Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · CVPR 2022 · arXiv:2112.10752 · version 2, 13 April 2022

arXiv 2112.107527/19/2026
Pixel-space diffusion repeats expensive work on a large grid; latent diffusion repeats it on a smaller representation.
Pixel-space diffusion repeats expensive work on a large grid; latent diffusion repeats it on a smaller representation.
AI / Technology / English

LoRA: Low-Rank Adaptation of Large Language Models

Authors: Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen Source: arXiv:2106.09685, version 2 (16 October 2021) · PDF Reading note: “Low rank” here describ

arXiv 2106.096857/19/2026
Full fine-tuning stores a full model per task; LoRA shares the base model and stores small task updates.
Full fine-tuning stores a full model per task; LoRA shares the base model and stores small task updates.
AI / Technology / English

Learning Transferable Visual Models From Natural Language Supervision

Alec Radford et al. · OpenAI · arXiv:2103.00020 · submitted 26 February 2021 Primary source: arXiv record · paper PDF

arXiv 2103.000207/19/2026
Comparison between a closed fixed-label system and natural-language supervision that can express a wider set of concepts.
Comparison between a closed fixed-label system and natural-language supervision that can express a wider set of concepts.
AI / Technology / English

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Authors: Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby Published: ICLR 2021; arXiv version 2, 3 June 2021 Primary source: [arX

arXiv 2010.119297/19/2026
A left-to-right view of an image becoming patches, tokens, a Transformer representation, and a prediction.
A left-to-right view of an image becoming patches, tokens, a Transformer representation, and a prediction.
AI / Technology / English

Denoising Diffusion Probabilistic Models

Jonathan Ho, Ajay Jain, Pieter Abbeel · NeurIPS 2020 · Paper · PDF

arXiv 2006.112397/19/2026
A clear image is gradually converted into Gaussian noise.
A clear image is gradually converted into Gaussian noise.
AI / Technology / English

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Patrick Lewis et al. · NeurIPS 2020 · arXiv:2005.11401 · version 4, 12 April 2021

arXiv 2005.114017/19/2026
A question is encoded, matched against a Wikipedia index, and combined with retrieved passages by a generator.
A question is encoded, matched against a Wikipedia index, and combined with retrieved passages by a generator.
AI / Technology / English

Language Models are Few-Shot Learners

Tom B. Brown et al. · OpenAI · 2020 · arXiv:2005.14165

arXiv 2005.141657/19/2026
Trending nowTwo-row comparison of task-specific fine-tuning and in-context learning
Two-row comparison of task-specific fine-tuning and in-context learning
AI / Technology / English

Scaling Laws for Neural Language Models

Jared Kaplan and colleagues, 2020 · arXiv:2001.08361

arXiv 2001.083617/19/2026
Three scaling inputs—model size, dataset size, and training compute—point toward lower test loss.
Three scaling inputs—model size, dataset size, and training compute—point toward lower test loss.
AI / Technology / English

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Authors: Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova Source: arXiv:1810.04805, version 2 · submitted 11 October 2018, revised 24 May 2019 Reading note: This explanation separates the authors’ reported findings from practical interp

arXiv 1810.048057/19/2026
Comparison of left-to-right context and BERT's jointly bidirectional context
Comparison of left-to-right context and BERT's jointly bidirectional context
AI / Technology / English

Attention Is All You Need

Authors: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin Published: arXiv:1706.03762, first submitted 12 June 2017; source version v7 dated 2 August 2023 Primary source: [arXiv abstract](https://arxiv.org/abs/17

arXiv 1706.037627/19/2026
Recurrence processes positions in a chain; self-attention connects and processes positions together during training.
Recurrence processes positions in a chain; self-attention connects and processes positions together during training.
AI / Technology / English

Deep Residual Learning for Image Recognition

Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015

arXiv 1512.033857/19/2026
A shallow network trains cleanly while a deeper plain stack becomes harder to optimize.
A shallow network trains cleanly while a deeper plain stack becomes harder to optimize.
AI / Technology / English

Auto-Encoding Variational Bayes

Diederik P. Kingma and Max Welling · arXiv:1312.6114 · submitted 20 December 2013, revised version available on arXiv Primary source: arXiv abstract · paper PDF

arXiv 1312.61147/19/2026
Why repeated inference becomes the bottleneck
Why repeated inference becomes the bottleneck
AI / Technology / English

Generative Adversarial Networks

Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014 · Primary source · PDF

arXiv 1406.26617/19/2026
Trending nowTwo paths feed the discriminator: generated samples arrive from noise through the generator, while real samples come from the dataset.
Two paths feed the discriminator: generated samples arrive from noise through the generator, while real samples come from the dataset.

FAQ

AI paper explainer FAQ

How is PaperBridge different from an AI paper summarizer?

A typical summarizer compresses the abstract and conclusion. PaperBridge also explains mechanisms, evidence quality, ablations, limitations, and practical meaning with original diagrams.

How can I find or submit a paper?

Search by title, paste an arXiv, paper-page, or public PDF link, or sign in and upload a public paper PDF. Existing reports open immediately; missing language versions can be generated.

Which languages are supported?

The public library currently focuses on Chinese, English, and Japanese reports. The generation workflow also supports Spanish.