Artem Zholus
I am a Member of Technical Staff at AMI Labs.
I am also a last year PhD student in MILA and
Polytechnique MontrΓ©al supervised by Prof. Sarath Chandar.
My ultimate research goal is to build adaptive and autonomous agents via World Models.
Previously I was a visiting researcher at FAIR @ Meta
advised by Mido Assran
(V-JEPA 2 and ???), and before that, a student researcher at
Google DeepMind
with Ross Goroshin
(TAPNext and TAPNext++).
Also, I had two internships at EPFL: at the
LIONS lab (in RL theory) under
Prof. Volkan Cevher and at VILAB
(in Multimodal Representation Learning) under Prof.
Amir Zamir. I obtained my Masters
degree at MIPT studying AI, ML, and Cognitive Sciences
and working at the CDS lab
under Prof. Aleksandr Panov
on task generalization in model-based reinforcement learning.
I received my BSc degree from ITMO University
majoring in Computer Science and Applied Mathematics.
|
|
News
- π’ August 2026 β Started as a Member of Technical Staff at
AMI Labs!
- π’ July 2026 β Finished my internship at
!
- π May 2026 β Awarded the prestigious Fonds de recherche du QuΓ©bec β Nature et technologies (FRQNT) Scholarship from the gouvernement du QuΓ©bec!
- π April 2026 β "TAPNext++" paper accepted to CVPR 2026 Findings!
- π June 2025 β V-JEPA 2 paper is out! Check out our blog post!
- π April 2025 β "TAPNext" paper accepted to ICCV 2025!
- π December 2024 β "BindGPT" paper accepted to AAAI 2025 with Best Poster Award!
- π’ October 2024 β Started an internship at
!
- π’ AprβSep 2024 β Internship at DeepMind
!
- π October 2023 β "Mastering Memory Tasks with World Models" paper accepted to ICLR 2024 with oral (top-1.2% of accepted papers)
- π May 2022 β "IGLU Gridworld" paper accepted to CVPR Workshop!
- π April 2022 β "IGLU 2022" paper accepted to NeurIPS Competition Track!
- π March 2022 β "Factorized World Models" paper accepted to ICLR Workshop!
|
|
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
Nilaksh*, Saurav Jha*, Artem Zholus*, Sarath Chandar
preprint, 2026
website /
arxiv /
Diffusion world modeling in semantic space preserves task semantics better than reconstruction spaces.
|
|
TAPNext++: What's Next for Tracking Any Point (TAP)?
Sebastian Jung*, Artem Zholus*, Martin Sundermeyer, Carl Doersch, Ross Goroshin, David Joseph Tan, Sarath Chandar, Rudolph Triebel, Federico Tombari
CVPR Findings, 2026
website /
arxiv /
code /
An extension of TAPNext to much longer video sequences, trained on 1024-frame sequences via parallelism, achieving state-of-the-art point tracking results for AR/XR and robotics.
|
|
Hierarchical Planning with Latent World Models
Wancong Zhang, Basile Terver, Artem Zholus, Soham Chitnis, Harsh Sutaria, Mido Assran, Randall Balestriero, Amir Bar, Adrien Bardes, Yann LeCun, Nicolas Ballas
arXiv preprint, 2026
website /
arxiv /
code /
A hierarchical model predictive control system that learns multi-scale world models in a unified latent space, letting long-horizon predictions guide short-horizon planning. Achieves strong zero-shot performance on real robot.
|
|
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran*, Adrien Bardes*, David Fan*, Quentin Garrido*, Russell Howes*, Mojtaba Komeili*, Matthew Muckley*, Ammar Rizvi*, Claire Roberts*, Koustuv Sinha*, Artem Zholus*, Sergio Arnaud*, Abha Gejji*, Ada Martin*, Francois Robert Hogan*, Daniel Dugas*, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier*, Yann LeCun*, Michael Rabbat*, and Nicolas Ballas*
Technical Report, 2025
website /
arxiv /
code /
blogpost /
hugginface /
By scaling world model pretraining to over a million hours of internet videos, we build V-JEPA 2 that excels at motion understanding, human-action anticipation, and video question answering. We show how action-conditioned post training on just 62 hours of unlabeled robot videos, enables zero-shot generalization in robotic control through planning in the latent space for tasks such as pick-and-place.
|
|
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
Artem Zholus, Carl Doersch, Yi Yang, Skanda Koppula, Viorica Patraucean, Xu Owen He, Ignacio Rocco, Mehdi S. M. Sajjadi, Sarath Chandar, Ross Goroshin
ICCV, 2025
website /
arxiv /
video /
code /
A new model for the Point Tracking task. Achieves SOTA with a huge margin while offering significantly faster online inference. We use a drastically different (from anything existing before for this task) approach, showing that only the scale of compute and data matters for this task.
|
|
BindGPT: A Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning
Artem Zholus, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Daniil Polykovskiy, Sarath Chandar, Alex Zhavoronkov
AAAI with Best Poster Award, 2025
website /
arxiv /
video /
hugginface /
BindGPT is a new framework for building drug discovery models that leverages compute-efficient pretraining, supervised funetuning, prompting, reinforcement learning, and tool use of LMs. This allows BindGPT to build a single pre-trained model that exhibits state-of-the-art performance in 3D Molecule Generation, 3D Conformer Generation, Pocket-Conditioned 3D Molecule Generation, posing them as downstream tasks for a pretrained model, while previous methods build task-specialized models without task transfer abilities.
|
|
Mastering Memory Tasks with World Models
Mohammad Reza Samsami*, Artem Zholus*, Janarthanan Rajendran, Sarath Chandar
ICLR with oral (top-1.2% of accepted papers), 2024
website /
arxiv /
openreview /
code /
The new State-of-the-Art performance in a diverse set of memory-intense Reinforcement Learning domains: bsuite (tabular, low dimensional), POPgym (tabular, high dimensional), Memory Maze (3D, embodied, high dimensional, long-term). Importantly, we reach super-human performance in Memory-Maze!
|
|
IGLU Gridworld: Simple and Fast Environment for Embodied Dialog Agents
Artem Zholus, Alexey Skrynnik, Shrestha Mohanty, Zoya Volovikova, Julia Kiseleva, Artur Szlam, Marc-Alexandre CotΓ©, Aleksandr I. Panov
Embodied AI workshop @ CVPR, 2022
arxiv /
code /
slides /
A lightweight reinforcement learning environment for building embodied agents with language context tasked to build 3D structures in Minecraft-like world.
|
|
IGLU 2022: Interactive Grounded Language Understanding in a Collaborative Environment at NeurIPS 2022
Julia Kiseleva*, Alexey Skrynni*, Artem Zholus*, Shrestha Mohanty*, Negar Arabzadeh*, Marc-Alexandre CΓ΄tΓ©*, Mohammad Aliannejadi, Milagro Teruel, Ziming Li, Mikhail Burtsev, Maartje ter Hoeve, Zoya Volovikova, Aleksandr Panov, Yuxuan Sun, Kavya Srinet, Arthur Szlam, Ahmed Awadallah
NeurIPS, Competition Track, 2022
website /
arxiv /
code /
AI competition where the goal is to follow a language instruction with context while being embodied in a 3D blocks world (RL track) and to ask a clarifying question in the case of ambiguity (NLP track).
|
|
Factorized World Models for Learning Causal Relationships
Artem Zholus, Yaroslav Ivchenkov, and Aleksandr Panov
OSC workshop, ICLR, 2022
arxiv /
code /
An RL agent that can generalize behavior on unseen tasks, which is done by learning a structured world model and constraining task specific information.
|
|