An agent can autonomously learn goal-oriented value functions that combine to solve new goal tasks specified as Boolean expressions, optimally and without further learning. This talk shows how these skills can also be composed over time, allowing the agent to satisfy complex temporal specifications, such as regular fragments of linear temporal logic, and achieve zero-shot transfer to unseen tasks.
Related to the paper Skill Machines: Temporal Logic Skill Composition in Reinforcement Learning.