Steve James
Steve James
Home
Publications
Research Lab
Contact
Workshop
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level …
Ebenezer Gelo
,
Geraud Nangue Tasse
,
Steven James
,
Benjamin Rosman
PDF
Cite
The Goal-Directed Frame for General Agents
Reinforcement learning is often framed around episodic, discounted, or average scalar rewards. While useful, these views miss a core …
Geraud Nangue Tasse
,
Steven James
,
Benjamin Rosman
PDF
Cite
Optimal Task Generalisation in Cooperative Multi-Agent Reinforcement Learning
While task generalisation is widely studied in the context of single-agent reinforcement learning (RL), little research exists in the …
Simon Rosen
,
Abdel Mfougouon Njupoun
,
Geraud Nangue Tasse
,
Steven James
,
Benjamin Rosman
PDF
Cite
ROSARL: Reward-Only Safe Reinforcement Learning
An important problem in reinforcement learning is designing agents that learn to solve tasks safely in an environment. A common …
Geraud Nangue Tasse
,
Tamlin Love
,
Mark Nemecek
,
Steven James
,
Benjamin Rosman
PDF
Cite
MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds
We propose a new benchmark for planning tasks based on the Minecraft game. Our benchmark contains 45 tasks overall, but also provides …
William Hill
,
Ireton Liu
,
Anita De Mello Koch
,
Damion Harvey
,
Nishanth Kumar
,
George Konidaris
,
Steven James
PDF
Cite
Counting Reward Automata: Sample Efficient Reinforcement Learning Through the Exploitation of Reward Function Structure
We present counting reward automata—a finite state machine variant capable of modelling any reward function expressible as a …
Tristan Bester
,
Benjamin Rosman
,
Steven James
,
Geraud Nangue Tasse
PDF
Cite
End-to-End Learning to Follow Language Instructions with Compositional Policies
We develop an end-to-end model for learning to follow language instructions with compositional policies. Our model combines large …
Vanya Cohen
,
Geraud Nangue Tasse
,
Nakul Gopalan
,
Steven James
,
Raymond Mooney
,
Benjamin Rosman
PDF
Cite
Skill Machines: Temporal Logic Composition in Reinforcement Learning
A major challenge in reinforcement learning is specifying tasks in a manner that is both interpretable and verifiable. One common …
Geraud Nangue Tasse
,
Devon Jarvis
,
Steven James
,
Benjamin Rosman
PDF
Cite
Augmentative Topology Agents For Open-ended Learning
In this work, we tackle the problem of open-ended learning by introducing a method that simultaneously evolves agents and increasingly …
Muhammad Nasir
,
Steven James
,
Christopher Cleghorn
PDF
Cite
Video
Supplementary Material
Facilitating Safe Sim-to-Real through Simulator Abstraction and Zero-shot Task Composition
Simulators are a fundamental part of training robots to solve complex control and navigation tasks. This is due to the speed and safety …
Tamlin Love
,
Devon Jarvis
,
Geraud Nangue Tasse
,
Branden Ingram
,
Steven James
,
Benjamin Rosman
PDF
Cite
Video
»
Cite
×