September 29, 2026

Behind the Music: How Spotify Built a Reinforcement Learning Simulator Using TensorFlow and TF-Agents

behind-the-music-how-spotify-built-a-reinforcement-learning-simulator-using-tensorflow-and-tf-agents

behind-the-music-how-spotify-built-a-reinforcement-learning-simulator-using-tensorflow-and-tf-agents

In the modern digital audio landscape, personal connection is everything. For millions of users worldwide, Spotify is not merely a library of songs; it is a dynamic companion that adapts to their morning commute, workout routine, and late-night study sessions. Behind this seamless auditory journey lies a complex web of machine learning algorithms designed to understand and predict human taste in real time.

Recently, a team of engineers and researchers at Spotify—comprising Surya Kanoria, Joseph Cauteruccio, Federico Tomasi, Kamil Ciosek, Matteo Rinaldi, and Zhenwen Dai—published groundbreaking work detailing how the streaming giant has tackled one of its most intricate challenges: sequential music recommendation. By leveraging TensorFlow and TF-Agents, the team successfully developed an offline simulation environment that mimics real user listening sessions, paving the way for advanced Reinforcement Learning (RL) applications at scale.


Main Facts: The Challenge of Sequential Recommendations

At its core, music recommendation is rarely a one-off decision. Instead, it is an ongoing, sequential process where algorithms must provide users with ordered sets of items—such as playlists or queues—that satisfy their shifting moods and intents over time.

Simulated Spotify Listening Experiences for Reinforcement Learning with TensorFlow and TF-Agents

Traditionally, machine learning models make predictions based on static user histories. However, these models often struggle to capture the evolving nature of a user’s listening session. If a user skips three upbeat tracks in a row, the recommendation engine needs to pivot instantly, adapting its strategy on the fly.

To solve this, Spotify turned to Reinforcement Learning, a machine learning paradigm where an "agent" learns to make decisions by interacting with an "environment" to maximize cumulative rewards. However, deploying an untrained RL agent directly into a production environment with real users poses severe risks. If an agent tries experimental or suboptimal strategies on live listeners, it could lead to frustration, churn, and a degraded user experience.

To bypass this roadblock, the Spotify team built a sophisticated, model-based offline simulator using TensorFlow and TF-Agents. This environment allowed them to prototype, train, and rigorously evaluate sequential recommendation agents safely away from live production servers.

Simulated Spotify Listening Experiences for Reinforcement Learning with TensorFlow and TF-Agents

Chronology: From Concept to Production-Scale RL

The journey toward deploying Reinforcement Learning across Spotify’s ecosystem required careful architectural planning and a step-by-step development process.

1. Aligning with the Production Stack

Early in the project, the team evaluated various RL frameworks. Because Spotify already relies heavily on TensorFlow and its extended ecosystem—including TFX and TensorFlow Serving—for its core production machine learning infrastructure, choosing TensorFlow Agents (TF-Agents) was a natural decision. This choice ensured that experimental models could eventually bridge the gap to production systems with maximum efficiency.

2. Designing the Offline Simulator

With the library chosen, the engineers faced a missing piece of the puzzle: a robust offline environment to simulate user behavior. Using TF-Agents’ core Environment primitives, the team designed a modular simulator capable of processing hypothetical user-content pairs. They integrated Keras-designed user models into the simulator to accurately predict how a simulated user would react to tracks recommended by an agent.

Simulated Spotify Listening Experiences for Reinforcement Learning with TensorFlow and TF-Agents

3. Developing Customized RL Agents

Standard single-choice slate recommendation algorithms were insufficient for Spotify’s needs, as users typically interact with multiple items across a listening session, and the available track pool is constantly changing. To address this high-dimensional state and action space, the team developed a novel Deep Q-Network variant called the Action-Head DQN (AH-DQN), alongside testing traditional algorithms like PPG and standard DQN.

4. Validation and Live Experimentation

Before rolling out the models to production, the team correlated their offline performance metrics with online results. Following successful offline trials, they launched live experiments, proving that their simulated performance estimates strongly aligned with real-world user engagement. The findings were later formalized and presented at KDD 2023.


Supporting Data: Inside the Architecture and Methodology

Building a realistic listening simulator requires translating complex human behaviors into structured mathematical components. The Spotify team structured their simulator around several key abstractions inspired by TF-Agents.

Simulated Spotify Listening Experiences for Reinforcement Learning with TensorFlow and TF-Agents

The RL Loop and Model-Based Simulation

In a standard RL loop, an agent observes the environment, takes an action based on its current policy, and receives a reward and a new observation from the environment. In Spotify’s formulation:

  • The State: The complete information summarizing the environment post-action.
  • The Observation: The specific subset of state information exposed to the agent.
  • The Reward: The simulated user’s predicted response (satisfaction metric) to the music recommendations driven by the agent’s action.

By relying on a serialized user model trained via Keras and unpacked within the simulator, the system calculates rewards safely without exposing live users to raw, untrained trial-and-error behaviors.

Modular Environment Abstraction

The team’s codebase utilizes a clean, abstract environment class that governs core behaviors:

Simulated Spotify Listening Experiences for Reinforcement Learning with TensorFlow and TF-Agents
class AbstractEnvironment(ABC):
    _user_model: AbstractUserModel = None
    _track_sampler: AbstractTrackSampler = None
    _episode_tracker: EpisodeTracker = None
    _episode_sampler: AbstractEpisodeSampler = None

    @abstractmethod
    def reset(self) -> List[float]:
      pass

    @abstractmethod
    def step(self, action: float) -> (List[float], float, bool):
      pass

    def observation_space(self) -> Dict:
      pass

    @abstractmethod
    def action_space(self) -> Dict:
      pass

Bridging Custom Code with TF-Agents

To leverage the powerful optimization and training loops native to TF-Agents, the team established a three-tier setup:

  1. Concrete Environment Classes: Implementing specific environments like PlaylistEnvironment that inherit from the abstract base class.
  2. Environment Builders: Using builder classes to ingest user models, track samplers, and episode samplers to instantiate the environment.
  3. TF-Agents Conversion: Wrapping the custom environment inside a TFAgtPyEnvironment subclass derived from py_environment.PyEnvironment, effectively converting it into a fully compliant TF-Agents environment ready for buffer storage and policy training.

Tackling Complexity with AH-DQN

To generate entire playlists from massive music catalogs, the Action-Head DQN evaluates the current state against each available action to produce a single Q-value. The algorithm iteratively selects the item with the highest Q-value, appends it to the slate, and repeats the process until the playlist is filled—neatly navigating the combinatorial complexity of sequential music recommendation.


Official Responses and Industry Recognition

The success of the project underscores the collaborative strength between industry engineering teams and open-source framework maintainers.

Simulated Spotify Listening Experiences for Reinforcement Learning with TensorFlow and TF-Agents

In their published findings, authors Federico Tomasi, Joseph Cauteruccio, Surya Kanoria, Kamil Ciosek, Matteo Rinaldi, and Zhenwen Dai expressed deep gratitude to their colleagues. They specifically acknowledged Mehdi Ben Ayed for early foundational work on the RL codebase and extended formal thanks to the TensorFlow Agents team for their continuous support, encouragement, and for maintaining the open-source library that made the entire architecture possible.

The research paper, titled "Simulated Spotify Listening Experiences for Reinforcement Learning with TensorFlow and TF-Agents," was officially accepted and presented at KDD 2023, garnering significant attention from the broader machine learning and recommender systems community.


Implications: The Future of Personalized Audio

The implications of Spotify’s successful offline simulator extend far beyond playlist generation. By proving that offline simulations can reliably predict online user behavior, Spotify has unlocked a scalable blueprint for applying Reinforcement Learning across numerous facets of its platform.

Simulated Spotify Listening Experiences for Reinforcement Learning with TensorFlow and TF-Agents

1. Risk-Free Experimentation

Machine learning researchers can now iterate rapidly on complex reward functions and policy updates in a controlled, offline environment. This drastically reduces the time-to-market for innovative recommendation features while protecting the end-user experience from erratic algorithmic behavior.

2. Deeper Personalization and Adaptability

As reinforcement learning models become more deeply embedded in audio streaming, applications will increasingly move away from static, genre-based categorization toward fluid, intent-driven curation. Whether a listener wants to discover underground electronic music or wind down with ambient sounds, the RL agent can dynamically curate sessions that honor microscopic shifts in user mood.

3. A Validation Template for the Industry

Spotify’s methodology provides a valuable case study for other large-scale recommendation platforms—spanning video streaming, e-commerce, and content publishing—facing similar sequential decision-making hurdles. By demonstrating how to seamlessly merge Keras user models, custom environment wrappers, and TF-Agents, Spotify has set a high architectural standard for production-grade reinforcement learning.