PhD Candidate @ UTN · Research Intern @ Ai2

Hi, I'm Tobi.

I build robots that learn. I'm a PhD candidate at the University of Technology Nuremberg, advised by Prof. Wolfram Burgard. My research is on embodied foundation models: making Vision-Language-Action models (VLAs) robust and reliable through reinforcement-learning-based fine-tuning, scalable synthetic data and sim-to-real transfer across robot embodiments.

Currently, I'm a PhD research intern with the robotics team at the Allen Institute for AI (Ai2) in Seattle, advised by Prof. Dieter Fox, working on zero-shot real-to-sim for VLAs and RL-based fine-tuning.

I also care about the infrastructure that makes robot learning reproducible. I lead Robot Control Stack, an open-source ecosystem for robot learning at scale.

  • Vision-Language-Action Models
  • Reinforcement Learning
  • Sim-to-Real
  • Cross-Embodiment
  • Robot Learning Infrastructure
Portrait of Tobias Jülg

Selected Publications

Papers

Full list on Scholar
RCS ecosystem: robots, sensors and applications
ICRA 2026 IEEE Int. Conf. on Robotics & Automation

Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale

Tobias Jülg, Pierre Krack, Seongjin Bien, Yannik Blei, Khaled Gamal, Ken Nakahara, Johannes Hechtl, Roberto Calandra, Wolfram Burgard, Florian Walter

VLAs and RL need lots of data and constant switching between simulation and real hardware, but classic robot software stacks were never designed for that workflow. RCS is a lean, ROS-free, pip-installable ecosystem with a layered C++/Python architecture. It exposes one Gymnasium interface for both MuJoCo digital twins and real robots, with teleoperation, data recording and VLA inference built in.

  • Supports four embodiments (Franka FR3, xArm7, UR5e, SO101), each with a digital twin
  • Evaluates Octo, OpenVLA and π0 across robots. Adding simulated data improves real-world VLA success
  • Records data at control rates of 90–120 Hz
@inproceedings{juelg2026robotcontrolstack,
  title={{Robot Control Stack}: {A} Lean Ecosystem for Robot Learning at Scale},
  author={Tobias J{\"u}lg and Pierre Krack and Seongjin Bien and Yannik Blei and Khaled Gamal and Ken Nakahara and Johannes Hechtl and Roberto Calandra and Wolfram Burgard and Florian Walter},
  booktitle={Proc.~of the IEEE Int.~Conf.~on Robotics \& Automation (ICRA)},
  year={2026}
}
RPD method: RL student guided by a VLA teacher
IROS 2025 IEEE/RSJ Int. Conf. on Intelligent Robots and Systems

Refined Policy Distillation: From VLA Generalists to RL Experts

Tobias Jülg, Wolfram Burgard, Florian Walter

Large VLAs generalize well but are slow and often not reliable enough on a specific task. RPD distills a VLA teacher (Octo or OpenVLA) into a compact RL expert. A PPO student is regularized toward the teacher's actions, so the VLA guides exploration while the student keeps improving through its own interaction.

  • Students outperform their VLA teachers on ManiSkill3 tasks, with dense and sparse rewards
  • Converge faster than plain PPO and stay robust to camera viewpoint changes
  • Fine-tuned Octo and OpenVLA checkpoints and the dataset are released on Hugging Face
@inproceedings{juelg2025refinedpolicydistillationvla,
  title={{Refined Policy Distillation}: {F}rom {VLA} Generalists to {RL} Experts},
  author={Tobias J{\"u}lg and Wolfram Burgard and Florian Walter},
  booktitle={Proc.~of the IEEE/RSJ Int.~Conf.~on Intelligent Robots and Systems (IROS)},
  year={2025}
}
RA-L 2026 IEEE Robotics and Automation Letters

Augmented Reality for RObots (ARRO): Pointing Visuomotor Policies Towards Visual Robustness

Reihaneh Mirjalili, Tobias Jülg, Florian Walter, Wolfram Burgard

Visuomotor policies break when the background, lighting or clutter changes. ARRO uses open-vocabulary segmentation to keep only the gripper and the task-relevant objects, and composites them onto a fixed virtual grid background. This happens in real time, during both training and inference, with no extra training or calibration.

  • Works with Diffusion Policy, Octo, OpenVLA and π0, in simulation and on a real robot
  • With distractors present, success rises from 30% (vanilla) to 90%
  • Enables real-to-sim and cross-embodiment transfer
@article{arro,
  author={Mirjalili, Reihaneh and J{\"u}lg, Tobias and Walter, Florian and Burgard, Wolfram},
  journal={IEEE Robotics and Automation Letters},
  title={Augmented Reality for RObots (ARRO): Pointing Visuomotor Policies Towards Visual Robustness},
  year={2026},
  volume={11},
  number={4},
  pages={4785-4792},
  doi={10.1109/LRA.2026.3665444}
}
FlowTouch: a robot imagines the tactile readings of a grasp before touching the object
IROS 2026 IEEE/RSJ Int. Conf. on Intelligent Robots and Systems

FlowTouch: View-Invariant Visuo-Tactile Prediction

Seongjin Bien, Carlo Kneissl, Tobias Jülg, Frank Fundel, Thomas Ressler-Antal, Florian Walter, Björn Ommer, Gitta Kutyniok, Wolfram Burgard

FlowTouch predicts what a vision-based tactile sensor will read before the robot touches anything. Instead of mapping camera images straight to touch, it reconstructs the object's 3D mesh with scene-reconstruction foundation models. A local point cloud around the planned contact then conditions a flow-matching generative model. Because it works from geometry, it is independent of the camera viewpoint and narrows the sim-to-real gap.

  • Predicted tactile images reach 86% grasp-stability accuracy, on par with real sensor readings
  • Reaches 81% zero-shot, without any samples from the target grasp dataset
  • Generalizes to a sensor instance never seen in training on a real FR3 with DIGIT sensors
@inproceedings{flowtouch,
  title={{FlowTouch}: View-Invariant Visuo-Tactile Prediction},
  author={Seongjin Bien and Carlo Kneissl and Tobias J{\"u}lg and Frank Fundel and Thomas Ressler-Antal and Florian Walter and Bj{\"o}rn Ommer and Gitta Kutyniok and Wolfram Burgard},
  booktitle={Proc.~of the IEEE/RSJ Int.~Conf.~on Intelligent Robots and Systems (IROS)},
  year={2026}
}

More publications

Architecture of the single-hidden-layer neural network: 27 inputs from 9 superlayers, 81 hidden nodes, outputs z and theta
NIM A 2025 Nuclear Instruments and Methods in Physics Research A

The Neural Network First-Level Hardware Track Trigger of the Belle II Experiment

S. Bähr, H. Bae, J. Becker, … , T. Jülg, … , K. Unger, J. Yin

Small neural networks running on FPGAs estimate each particle track's origin along the beam within the microsecond budget of Belle II's first-level trigger. This lets the trigger reject background events in hardware while keeping physics events.

@article{baehr2025neural,
  title={The neural network first-level hardware track trigger of the {Belle II} experiment},
  author={B{\"a}hr, S. and Bae, H. and Becker, J. and Bertemes, M. and Campajola, M. and Ferber, T. and Forsthofer, T. and Hiesl, S. and Inguglia, G. and Iwasaki, Y. and J{\"u}lg, T. and Kiesling, C. and Knoll, A. C. and Koga, T. and Lai, Y.-T. and Lenz, A. and Liu, Y. and Meggendorfer, F. and Nakazawa, H. and Neu, M. and Schieck, J. and Schmidt, E. and Shiu, J.-G. and Skambraks, S. and Unger, K. and Yin, J.},
  journal={Nuclear Instruments and Methods in Physics Research Section A},
  volume={1073},
  pages={170279},
  year={2025},
  doi={10.1016/j.nima.2025.170279}
}
SOM-AE architecture: an autoencoder compresses images into an 8-dimensional latent that feeds a multi-modal self-organizing map
AMAM 2023 Int. Symposium on Adaptive Motion of Animals and Machines

Multi-Modal Representation Learning for Mapping Between Body Motion and Visual Imagery

Tobias Jülg, Florian Walter, Dongmin Kim, Hoshinori Kanazawa, Alois Knoll, Yasuo Kuniyoshi

A multi-modal self-organizing map learns how a simulated infant's arm movements relate to what it sees. Compressing the images with an autoencoder first shrinks the map from 10.5 GB to 33 MB and cuts training from 61 h to 15 min, with equal or better mapping quality.

@inproceedings{juelg2023multimodal,
  title={Multi-Modal Representation Learning for Mapping Between Body Motion and Visual Imagery},
  author={J{\"u}lg, Tobias and Walter, Florian and Kim, Dongmin and Kanazawa, Hoshinori and Knoll, Alois and Kuniyoshi, Yasuo},
  booktitle={The 11th International Symposium on Adaptive Motion of Animals and Machines (AMAM)},
  pages={150--151},
  year={2023},
  doi={10.18910/92312}
}

Open Source

Software

More on GitHub
VLAgents architecture: environments talk to a policy server over shared memory or TCP

VLAgents

A policy server for efficient VLA inference. Policies such as Octo, OpenVLA, π0 and LeRobot sit behind one Gymnasium-style protocol. It uses zero-copy shared memory in simulation and compressed TCP streaming for remote robots.

FrankIK

A fast analytical inverse kinematics solver for Franka Panda and FR3, written in C++ with Python bindings. Each query takes about 6.5 µs, roughly 2× faster than the fastest Python IK library.

robotiq2f

A pure-Python, ROS-free driver for the Robotiq 2F-85 gripper over Modbus RTU. It finds the device by serial number (so it still works after replugging), with non-blocking status polling and TCP-offset compensation.

Open Data

Datasets

More on Hugging Face
Head-camera frames from eight DuoBench tasks, real and simulated Real Transfer-Cube episode from head and both wrist cameras

DuoBench FR3 Duo

Teleoperated bimanual demonstrations on a dual Franka Research 3. It covers 11 simulated and 4 real tasks, with a language instruction per task, in LeRobot format ready for training.

  • 750 episodes
  • 446k frames
  • 30 fps
  • 3 cameras
  • LeRobot v3
  • 4.5 GB
Third-person and wrist camera frames of a Franka picking a green box One teleoperated pick episode, third-person and wrist camera

Green Box Pick (FR3)

Real-world teleoperated Franka FR3 demonstrations of the task "pick the green box", recorded with RCS. Used to fine-tune Octo, OpenVLA and π0 in the RCS paper.

  • 143 episodes
  • 36.7k frames
  • 10 fps
  • 2 cameras
  • LeRobot v2.1
  • 5.3 GB
Frames from six ManiSkill3 tasks: PickCube, PushCube, PullCube, LiftPegUpright, RollBall, PushT One PushT episode from both cameras

ManiSkill3 in RLDS

ManiSkill3 demonstrations of a simulated Franka Panda, converted to RLDS so VLAs like Octo and OpenVLA can be fine-tuned on them. Used for the RPD teachers.

  • 7,993 episodes
  • Pick · Push · Pull · Lift · Roll · PushT
  • RLDS / TFDS
  • 45 GB

Background

Positions & Education

Research & Industry

  1. 06/2026 – now PhD Research Intern Allen Institute for AI (Ai2) · Robotics, Seattle Advisor: Dieter Fox
  2. 10/2023 – now Doctoral Researcher UTN · Lab for AI & Robotics, Nuremberg Advisor: Wolfram Burgard
  3. 12/2021 – 09/2022 Machine Learning Intern Max Planck Institute of Physics · Belle II, Munich
  4. 05/2021 – 07/2022 Deep Learning Intern Cruise · Radar perception, Munich

Education

  1. 10/2023 – now Ph.D. in Robotics University of Technology Nuremberg (UTN)
  2. 10/2022 – 03/2023 Visiting Research Graduate The University of Tokyo · Intelligent Systems and Informatics Lab Master's thesis
  3. 10/2020 – 09/2023 M.Sc. Informatics Technical University of Munich (TUM) Grade 1.0, high distinction (top 2%)
  4. 10/2016 – 05/2020 B.Sc. Informatics Technical University of Munich (TUM)