Behind the Music: How Spotify Built Simulated Reinforcement Learning Environments Using TensorFlow and TF-Agents

In the modern era of digital audio streaming, personalization is the invisible engine that drives user satisfaction. When millions of listeners open Spotify every day, they expect a seamless, intuitive journey through an infinite catalog of songs, podcasts, and playlists. Meeting this expectation is not merely a matter of suggesting music that fits a user’s perennial tastes; it requires a deep, real-time understanding of evolving intent, mood, and context as a listening session unfolds.
To tackle this complex challenge, a team of researchers and engineers at Spotify—Surya Kanoria, Joseph Cauteruccio, Federico Tomasi, Kamil Ciosek, Matteo Rinaldi, and Zhenwen Dai—turned to Reinforcement Learning (RL). By viewing music recommendation as a sequential decision-making process, the team sought to craft intelligent agents capable of adapting to users over time. However, training reinforcement learning agents directly on live production systems poses severe risks, chief among them the potential degradation of user experience through poorly calibrated, exploratory recommendations.
To bridge the gap between theoretical machine learning and safe production deployment, the Spotify team engineered a robust offline simulation environment using TensorFlow and TF-Agents. Their breakthrough, detailed extensively at KDD 2023, marks a significant milestone in scaling reinforcement learning applications across the world’s leading audio platform.
The Sequential Nature of Music Discovery
To understand why Spotify embraced reinforcement learning, one must first examine the inherent limitations of traditional static recommendation systems. In standard collaborative filtering or content-based recommendation frameworks, the system evaluates a user’s historical preferences and outputs a fixed set of items. While effective for single-point recommendations, these models often falter when faced with multi-item sequences, such as a continuous queue of tracks or an automatically generated playlist.

Music listening is a dynamic, chronological journey. A user’s reaction to the first song in a playlist directly influences their tolerance, appetite, and mood for the subsequent tracks. If a user skips three upbeat pop songs in a row, the system must rapidly pivot to a different tempo or genre. This represents a classic sequential decision-making problem where actions taken at time step $t$ alter the state of the environment and influence cumulative future rewards.
Reinforcement learning is mathematically designed for such environments. In an ideal RL loop, an agent observes the state of an environment, selects an action (such as recommending a specific song or sequence of songs), receives a reward based on the user’s reaction (such as a stream completion, a like, or a skip), and updates its internal policy to maximize long-term satisfaction.
Yet, deploying an untrained or semi-trained RL agent directly into a live production environment with millions of real-world users is out of the question. Unconstrained exploration could lead to repetitive, jarring, or outright frustrating music recommendations, causing users to abandon the application. To safely iterate, test, and train these agents, Spotify needed a high-fidelity offline simulator.
Chronology of the Project: From Concept to Production
The development of Spotify’s simulated listening experience framework followed a deliberate, multi-phase technological roadmap.

Phase 1: Architectural Alignment and Library Selection
Early in the exploratory phase, the engineering team evaluated various machine learning ecosystems to support their RL ambitions. Because Spotify’s existing production machine learning stack relied heavily on TensorFlow and its extended ecosystem—including TensorFlow Extended (TFX) and TensorFlow Serving—the team made the strategic decision to adopt TensorFlow Agents (TF-Agents).
Choosing TF-Agents ensured that experimental models developed in research notebooks could seamlessly integrate into Spotify’s existing infrastructure down the line. Furthermore, the modular design primitives native to TF-Agents provided a natural blueprint for constructing an offline Spotify simulator.
Phase 2: Building the User Model and Simulator Core
Without live users to interact with during early training, the team required a reliable proxy. They adopted a model-based RL approach, leveraging Keras to design and train sophisticated user models. These models were trained on historical listening data to predict how a hypothetical user would respond to a given sequence of tracks.
Once trained, these user models were serialized and unpacked by a custom offline simulator. Built on top of TF-Agents environment primitives, the simulator was designed to replicate real-world listening sessions. It combined several modular components:

- Episode Samplers (
_episode_sampler): Responsible for initializing simulations with representations of hypothetical users and content pools derived from real-world listening data. - Track Samplers (
_track_sampler): Solved the combinatorial complexity of slate recommendations by translating the agent’s abstract continuous or discrete actions into concrete, ordered lists of track recommendations. - Episode Trackers (
_episode_tracker): Expanded beyond standard TF-Agents replay buffers to log granular user model predictions, sampled user-content pairs, and specialized diagnostic metrics for debugging.
Phase 3: Developing Customized Agents (AH-DQN)
Standard slate recommendation algorithms typically focus on selecting a single item or a small, independent set of items. However, Spotify’s playlist generation task required recommending an entire ordered sequence of tracks where user feedback occurs across multiple items, and the underlying track pool is constantly changing.
To handle the massive state and action spaces inherent to this problem, the team developed a novel variant of the Deep Q-Network known as the Action-Head DQN (AH-DQN). Instead of outputting a single Q-value for an entire collection, the AH-DQN evaluates the current state against individual available actions iteratively, identifying the highest-value track to add to the playlist slate until the queue is complete.
Phase 4: Offline Validation and Large-Scale Online Deployment
Before rolling out the models to production, the team conducted rigorous offline evaluations to test various policies (naive, heuristic, model-driven, and RL-based). Crucially, they compared these simulated offline performance estimates against real-world online A/B testing results.
The findings were definitive: offline performance estimates generated by the TF-Agents simulator exhibited strong directional correlation with live online metrics. This validation removed the primary bottleneck for reinforcement learning research at Spotify, clearing the path for large-scale, automated experimentation.

Technical Architecture: Code and Components
At the heart of the Spotify simulator is a clean, abstract environment architecture that mirrors the structure of standard TF-Agents environments. The core abstraction requires concrete instantiations to define foundational methods for environment resets, steps, observation spaces, and action spaces:
class AbstractEnvironment(ABC):
_user_model: AbstractUserModel = None
_track_sampler: AbstractTrackSampler = None
_episode_tracker: EpisodeTracker = None
_episode_sampler: AbstractEpisodeSampler = None
@abstractmethod
def reset(self) -> List[float]:
pass
@abstractmethod
def step(self, action: float) -> (List[float], float, bool):
pass
def observation_space(self) -> Dict:
pass
@abstractmethod
def action_space(self) -> Dict:
pass
To transition this custom Python environment into a fully functional training ground compatible with TF-Agents, the Spotify team engineered a multi-step integration pipeline. First, concrete environments—such as the PlaylistEnvironment—are instantiated using specific user models and track samplers via an Environment Builder Class.
Second, a conversion wrapper (TFAgtPyEnvironment) subclasses TF-Agents’ py_environment.PyEnvironment, wrapping the custom Spotify environment:
class TFAgtPyEnvironment(py_environment.PyEnvironment):
def __init__(self, environment: AbstractEnvironment):
super().__init__()
self.env = environment
Finally, the environment builder compiles this wrapper into a native TensorFlow environment (tf_env), making it immediately available for training algorithms like PPG, standard DQN, and the custom AH-DQN.

Supporting Data and Empirical Results
The success of any simulation-driven machine learning initiative rests on the fidelity of its proxies. In their KDD 2023 paper, the Spotify research team published comparative analyses evaluating how well simulated offline performance metrics tracked real-world user engagement.
- Correlation Accuracy: Across numerous naive, heuristic, model-driven, and reinforcement learning policies, the team observed a powerful directional alignment between simulated performance scores and scaled online rewards.
- Session Termination Modeling: Simulation termination dynamics were finely tuned using empirical data. For instance, empirical investigations revealed precise session drop-off curves—such as determining that 92% of specific listening sessions terminate after precisely six sequential track skips—allowing the simulator to mimic realistic user fatigue and disengagement.
- Scalability: By shifting the heavy lifting of agent training to offline TensorFlow-backed environments, the team drastically reduced compute expenditure and eliminated the risk of exposing active users to unvetted exploratory policies.
Official Responses and Perspectives
The publication of the KDD 2023 paper—authored by Federico Tomasi, Joseph Cauteruccio, Surya Kanoria, Kamil Ciosek, Matteo Rinaldi, and Zhenwen Dai—has drawn widespread praise from both the academic reinforcement learning community and industrial machine learning practitioners.
In their published acknowledgements, the authors extended special gratitude to early contributors such as Mehdi Ben Ayed for laying the groundwork for the initial RL codebase. Furthermore, the team publicly thanked the TensorFlow Agents open-source development team for providing the flexible, robust library architecture that made the entire end-to-end simulation pipeline possible.
Industry analysts have noted that Spotify’s methodology provides a blueprint for other recommendation-heavy platforms—such as video streaming, e-commerce, and news aggregators—that struggle with the tension between exploration-based learning and user retention.

Broader Implications for Audio Streaming and Machine Learning
The implications of Spotify’s successful integration of TF-Agents and custom simulation environments extend far beyond automated playlist generation.
- Democratizing Reinforcement Learning in Production: By proving that offline simulators can accurately predict online performance, Spotify has lowered the barrier to entry for complex RL architectures in production environments. Engineering teams no longer need to gamble user satisfaction on raw online exploration.
- Handling Combinatorial Slate Complexity: The invention and deployment of the Action-Head DQN (AH-DQN) provides a scalable template for recommending ordered collections of items where multiple interaction points exist per session. This architecture can readily be adapted to podcast chapters, audiobook queues, and multi-format media feeds.
- Ecosystem Synergy: The project underscores the immense value of adhering to established machine learning ecosystems. By keeping their stack anchored in TensorFlow, TFX, and TF-Agents, Spotify ensured that research innovations could transition smoothly into industrial-scale production engines without requiring costly rewrites or infrastructure overhauls.
As Spotify continues to innovate in the audio streaming space, the marriage of reinforcement learning and high-fidelity simulation ensures that the music will keep playing—smarter, smoother, and more attuned to the human experience than ever before.
