Reward Functions That Evolve with Agent Learning
1. Defining Reward Functions and Their Role in Agent Learning
Defining Reward Functions and Their Role in Agent Learning
Reward functions serve as the foundational mechanism for shaping an agent's behavior in reinforcement learning (RL). Mathematically, a reward function R maps a state-action pair (s, a) or a state transition (s, a, s') to a scalar value r ∈ ℝ, which quantifies the desirability of the agent's action in a given state. The agent's objective is to maximize the expected cumulative reward, often formalized as the return Gt:
where γ ∈ [0, 1] is the discount factor, trading off immediate versus long-term rewards. The reward function is not merely a scoring mechanism; it encodes the task's objective and implicitly defines the agent's policy π(a|s) through optimization. Poorly designed rewards can lead to unintended behaviors, such as reward hacking, where the agent exploits loopholes to maximize rewards without achieving the intended goal.
Properties of Effective Reward Functions
An ideal reward function balances three critical properties:
- Sparse vs. Dense Rewards: Sparse rewards (e.g., +1 upon task completion) simplify design but hinder exploration. Dense rewards (e.g., incremental penalties/progress signals) guide learning but risk misalignment with the true objective.
- Credit Assignment: The function must attribute rewards to actions causally responsible for outcomes, especially in delayed-reward scenarios. Temporal difference methods like TD(λ) address this by propagating rewards backward.
- Scale and Normalization: Rewards must be scaled to avoid gradient instability. Techniques like reward clipping or adaptive normalization (e.g., Pop-Art) maintain stable learning dynamics.
Dynamic Reward Formulations
Static reward functions often fail in complex environments. Dynamic alternatives adapt to the agent's proficiency or environmental changes:
Here, β(t) is a time-dependent weighting factor, and Raux is an auxiliary reward that may decay as the agent masters sub-tasks. For example, in robotic manipulation, a distance-based reward for reaching a target might phase out as the agent learns precise grasping.
Case Study: Inverse Reinforcement Learning
When explicit reward design is infeasible, inverse RL (IRL) infers R(s, a) from expert demonstrations. The MaxEnt IRL framework models the expert's policy as Boltzmann-rational:
where τ is a trajectory. This approach reveals how reward functions can be learned rather than handcrafted, bridging the gap between human intent and agent behavior.
1.2 Static vs. Dynamic Reward Functions: Key Differences
Definition and Core Characteristics
A static reward function remains fixed throughout the agent's learning process, providing a constant mapping from states and actions to scalar rewards. In contrast, a dynamic reward function adapts based on the agent's performance, environmental changes, or external feedback. The mathematical formulation for a static reward function is straightforward:
where s represents the state, a the action, and f a predefined function. Dynamic reward functions introduce time-dependence or state-history dependence:
Here, \(\mathcal{H}_t\) denotes the history of states, actions, and rewards up to time t.
Behavioral Implications in Reinforcement Learning
Static reward functions are predictable but may lead to suboptimal policies if the environment is non-stationary. For example, in robotic control, a fixed reward for reaching a target position may fail if the target moves. Dynamic reward functions can address this by incorporating real-time feedback, such as:
- Reward shaping based on the agent's recent performance
- Adaptive penalties for repeated failures
- Curriculum learning, where rewards evolve to guide the agent from simple to complex tasks
Mathematical Adaptability
Dynamic reward functions often employ meta-learning techniques. Consider a reward function that adapts via gradient descent:
where \(\alpha\) is a learning rate, \(\mathcal{L}\) is a loss function comparing the optimal policy \(\pi^*\) to the current policy \(\pi_t\). This approach is common in inverse reinforcement learning.
Practical Applications and Challenges
In autonomous driving, static rewards for lane-keeping may suffice for highway scenarios, but dynamic rewards are essential for urban environments with pedestrians and unpredictable traffic. Key challenges include:
- Non-stationarity: The agent must distinguish between environmental changes and reward function adaptations.
- Credit assignment: Determining which actions led to reward changes becomes complex.
- Convergence: Theoretical guarantees for policy convergence are harder to establish.
Case Study: Multi-Agent Systems
In competitive multi-agent settings like StarCraft II, static rewards for resource collection can lead to exploitable strategies. Dynamic rewards that penalize over-reliance on a single strategy promote robust adaptation. The reward function might incorporate an entropy term:
where H is the policy entropy and \(\beta\) controls exploration incentives.
1.3 Challenges in Designing Effective Reward Functions
Reward Hacking and Specification Gaming
A fundamental challenge in reinforcement learning (RL) is reward hacking, where an agent exploits loopholes in the reward function to maximize returns without achieving the intended objective. This occurs when the reward function fails to fully capture the desired behavior, leading to degenerate solutions. For example, an RL agent trained to maximize game scores might discover unintended strategies that inflate scores without meaningful progress.
The mathematical formulation of this problem can be expressed as:
where the optimal policy \(\pi^*\) maximizes cumulative reward but may do so in ways misaligned with designer intent. This discrepancy between reward function and true objective is known as the specification problem.
Sparse and Delayed Rewards
In many real-world tasks, rewards are sparse (only given at task completion) or delayed (consequences of actions manifest much later). This creates a credit assignment problem where the agent struggles to associate actions with long-term outcomes. The temporal difference error \(\delta_t\):
becomes noisy when \(r_t\) is zero for most timesteps, leading to unstable learning. Environments like robotic control or strategic games often exhibit this property, requiring advanced techniques like reward shaping or hierarchical RL.
Non-Stationarity in Multi-Agent Systems
When multiple learning agents interact, the reward function becomes non-stationary as other agents' policies evolve. Consider a two-player zero-sum game with Q-learning updates:
Each agent's update affects the other's reward landscape, potentially leading to oscillating or divergent behaviors. This is particularly problematic in competitive environments like auctions or adversarial scenarios.
Scalability and Curse of Dimensionality
Designing reward functions that scale to high-dimensional state spaces requires careful feature engineering. The curse of dimensionality manifests when the state space \(\mathcal{S}\) grows exponentially with system complexity:
Manual reward shaping becomes infeasible, necessitating automatic reward learning methods like inverse reinforcement learning (IRL) or preference-based learning. However, these approaches introduce their own challenges in terms of sample efficiency and human feedback reliability.
Ethical and Safety Considerations
Poorly designed reward functions can lead to unsafe or unethical behaviors. The side effects problem occurs when maximizing rewards causes undesirable environmental changes, while the reward tampering problem involves agents manipulating their reward signal. Formal frameworks like constrained RL:
attempt to mitigate these risks but require careful constraint specification and verification.
Partial Observability and Noisy Rewards
In partially observable Markov decision processes (POMDPs), the agent receives rewards based on incomplete state information. The belief state \(b(s)\) represents the probability distribution over true states:
where \(\eta\) is a normalizing constant. Noisy or misleading rewards in such environments can cause the agent to learn incorrect correlations between observations and desirable outcomes.
2. Reward Shaping and Its Impact on Learning Efficiency
Reward Shaping and Its Impact on Learning Efficiency
Reward shaping introduces auxiliary rewards to guide an agent toward desired behaviors more efficiently than sparse environmental rewards alone. The shaped reward function \( R'(s, a, s') \) augments the original reward \( R(s, a, s') \) with a potential-based term \( \Phi(s') - \Phi(s) \), where \( \Phi \) is a potential function encoding domain knowledge:
This formulation preserves policy invariance when \( \gamma \) matches the discount factor of the MDP, as proven by Ng et al. (1999). The potential function \( \Phi \) acts as a differentiable heuristic, providing intermediate learning signals that mitigate credit assignment problems in long-horizon tasks.
Gradient-Based Potential Functions
Modern implementations often learn \( \Phi \) simultaneously with the policy using gradient descent. The potential function's parameters \( \theta_\Phi \) are updated to minimize the temporal difference error:
This creates a symbiotic relationship where \( \Phi \) provides denser rewards for policy learning while being refined by the agent's experience. In deep RL, \( \Phi \) is typically implemented as a neural network with fewer layers than the policy network to prevent overfitting.
Impact on Sample Efficiency
Reward shaping alters the optimization landscape in three key ways:
- Curriculum Effect: The shaped reward creates a smoother gradient towards optimal policies, particularly useful in environments with deceptive local optima
- Credit Assignment: Shortens the temporal distance between actions and consequences in sparse reward settings
- Exploration Guidance: The potential function's gradient \( \nabla \Phi(s) \) provides directional hints for exploration strategies
Empirical studies show sample efficiency improvements of 3-10x in benchmark tasks like Montezuma's Revenge when combining potential-based shaping with intrinsic motivation. However, poorly designed shaping rewards can lead to reward hacking, where the agent exploits the shaping function without solving the intended task.
Dynamic Shaping Strategies
Advanced implementations adapt the shaping magnitude during training:
where \( \tau \) controls the decay rate of shaping influence. This annealing schedule allows the agent to gradually transition from shaped rewards to the true environmental rewards, preventing over-reliance on the shaping function.
Recent work in meta-learning has extended this concept by parameterizing the shaping function as \( \Phi_\omega(s) \), where \( \omega \) is adapted online using gradient-based meta-optimization. This enables the shaping function to evolve at a different timescale than the policy itself.

2.3 Dynamic Reward Adjustment Based on Agent Performance
Dynamic reward adjustment mechanisms modify the reward function in response to the agent's learning progress, ensuring that the feedback remains relevant as the agent's policy evolves. This approach addresses the challenge of reward sparsity or misalignment that arises when a fixed reward function fails to guide the agent effectively beyond initial exploration phases.
Performance-Based Reward Shaping
One common method involves scaling rewards based on the agent's recent performance. Let Rt be the original reward at time t, and let μk and σk represent the mean and standard deviation of rewards over a sliding window of the last k episodes. The adjusted reward R't can be computed as:
where ϵ is a small constant for numerical stability. This normalization prevents reward magnitudes from becoming too large or too small as the agent improves, maintaining stable gradient updates.
Curriculum Learning via Reward Adjustment
In curriculum learning, the reward function is progressively modified to guide the agent from simpler to more complex behaviors. For instance, in a navigation task, the reward for reaching intermediate waypoints might be initially high but decay as the agent masters those sub-tasks, shifting emphasis toward the final goal. The decay can be modeled as:
where λ controls the decay rate, and success_rate measures the agent's recent performance on the sub-task.
Potential-Based Reward Shaping
Potential-based reward shaping ensures that dynamic adjustments do not alter the optimal policy. Given a potential function Φ(s) that encodes desirable states, the shaped reward R't is:
where γ is the discount factor. The potential function can be updated periodically based on the agent's current capabilities, such as using the value function Vπ(s) of the current policy π.
Practical Implementation Considerations
- Sliding Window Size: A small window reacts quickly to performance changes but may overfit noise, while a large window provides stability at the cost of adaptability.
- Decay Scheduling: The decay rate λ must balance between retaining useful guidance and avoiding premature reward sparsity.
- Convergence Monitoring: Dynamic rewards can introduce non-stationarity, so tracking policy convergence or divergence is critical.
These techniques are widely applied in robotics, game AI, and autonomous systems where the agent's objectives may scale in complexity or shift over time.

3. Genetic Algorithms for Reward Function Optimization
3.1 Genetic Algorithms for Reward Function Optimization
Genetic algorithms (GAs) provide a biologically inspired optimization framework for evolving reward functions in reinforcement learning (RL). By treating reward functions as genotypes subject to selection, crossover, and mutation, GAs enable adaptive reward shaping that co-evolves with the agent's policy. This approach is particularly effective in sparse-reward or deceptive environments where hand-designed reward functions fail to guide learning.
Mathematical Framework
The GA operates on a population of reward function candidates Ri, each represented as a parameterized function:
where θ are the evolvable parameters and φ is a state-action-next-state feature mapping. The fitness Fi of each reward function is evaluated by training an RL agent with Ri for N episodes and measuring the agent's final performance:
where π*Ri is the optimal policy under reward function Ri.
Genetic Operators for Reward Functions
The evolutionary process applies three key operators:
- Selection: Tournament selection with size k samples k reward functions and selects the one with highest fitness
- Crossover: Two-point crossover swaps segments of the parameter vectors θ between parent reward functions
- Mutation: Gaussian noise added to θ with probability pm and variance σ2
The mutation operator is particularly crucial for maintaining exploration in the reward function space. Adaptive mutation rates that decrease with generation number often outperform fixed rates:
where g is the generation number and λ controls the decay rate.
Practical Implementation Considerations
When implementing GA-based reward optimization, several architectural choices significantly impact performance:
- Parallel Evaluation: Distributed fitness evaluation across multiple agents or environments dramatically reduces wall-clock time
- Elitism: Preserving the top k reward functions unchanged between generations prevents regression
- Diversity Maintenance: Techniques like fitness sharing or novelty search prevent premature convergence
The reward function representation must balance expressiveness with evolvability. Neural networks with 1-2 hidden layers typically outperform both linear functions (too constrained) and deep networks (too difficult to evolve).
Case Study: Maze Navigation
In a sparse-reward maze environment where the agent only receives reward upon reaching the goal, a GA-evolved reward function discovered that rewarding:
where g is the goal state and s0 is the initial state, led to 3.2× faster convergence than sparse rewards alone. The evolved function implicitly shaped the reward to guide the agent toward the goal while maintaining the global optimum.
Convergence Properties
The GA's convergence can be analyzed using the schema theorem. For a reward function schema H with defining length δ(H) and order o(H), the expected number of instances in the next generation is:
where f(H) is the average fitness of instances of H, favg is the population average fitness, pc is the crossover probability, pm is the mutation probability, and L is the chromosome length. This shows how building blocks of high-fitness reward functions proliferate through generations.
Co-Evolution of Agents and Reward Functions
The co-evolution of agents and reward functions represents a paradigm where the reward function is not static but dynamically adapts alongside the learning agent. This approach addresses limitations in traditional reinforcement learning (RL), where fixed reward functions may lead to suboptimal behaviors or reward hacking—scenarios where agents exploit loopholes in the reward structure without achieving the intended objective.
Mathematical Framework
Consider a Markov Decision Process (MDP) defined by the tuple (S, A, P, R, γ), where S is the state space, A the action space, P the transition dynamics, R the reward function, and γ the discount factor. In co-evolutionary RL, the reward function R is parameterized by a set of learnable parameters θ, making it Rθ(s, a, s'). The agent’s policy πφ, parameterized by φ, and the reward function Rθ are optimized jointly.
This joint optimization can be framed as a bilevel problem: the inner loop optimizes the agent’s policy given a reward function, while the outer loop adjusts the reward function to guide the agent toward desired behaviors.
Mechanisms for Co-Evolution
Several mechanisms enable the co-evolution of agents and reward functions:
- Inverse Reinforcement Learning (IRL): The reward function is inferred from expert demonstrations, then refined as the agent learns.
- Reward Shaping: The reward function is adjusted dynamically to provide intermediate feedback, often using potential-based methods to preserve policy invariance.
- Meta-Learning: The reward function is optimized via meta-RL, where the outer loop learns a reward function that maximizes the agent’s performance across tasks.
Practical Challenges
Co-evolution introduces complexities such as:
- Non-Stationarity: The agent’s policy and reward function change simultaneously, leading to a non-stationary learning environment.
- Credit Assignment: Determining whether poor performance stems from the agent’s policy or the reward function’s design becomes ambiguous.
- Convergence: Joint optimization may lead to unstable training dynamics, requiring careful balancing of learning rates for θ and φ.
Case Study: Self-Adversarial Reward Learning
In self-adversarial frameworks, the reward function is trained to discriminate between the agent’s behaviors and desired behaviors, akin to Generative Adversarial Networks (GANs). The discriminator (reward function) and generator (policy) are co-optimized:
where π* represents expert demonstrations. This approach has been applied successfully in robotics and game-playing agents.
Applications
Co-evolutionary reward design is particularly useful in:
- Robotics: Adapting reward functions for complex manipulation tasks where hand-designed rewards are infeasible.
- Game AI: Dynamically balancing difficulty by adjusting rewards based on player skill.
- Autonomous Systems: Ensuring safety by evolving reward functions that penalize risky behaviors.
3.3 Case Studies: Successes and Limitations
AlphaGo's Adaptive Reward Shaping
The AlphaGo system demonstrated how reward functions can evolve through multiple training phases. Initially, the agent learned from human expert games with a simple win/loss reward R0(s) = ±1. During policy iteration, the reward function incorporated a value network's predictions:
where λ controlled the blending between immediate and long-term rewards. This approach succeeded because:
- The value network Vθ provided denser rewards than sparse win/loss signals
- Gradual adaptation prevented reward hacking behaviors
- The final policy achieved superhuman performance despite imperfect reward proxies
Robotics: Curriculum Learning Failures
In robotic arm manipulation tasks, dynamically adjusted rewards based on task progress often led to suboptimal policies. A 2021 study on door-opening tasks revealed:
where reward adaptation rate α was tied to policy gradient magnitudes. Key limitations included:
- Premature convergence when rewards adapted too quickly
- Non-stationarity making value estimation unstable
- Sensitivity to hyperparameters in the adaptation rule
Multi-Agent Systems: Emergent Reward Dynamics
The Hanabi Challenge showed how co-adapting reward functions across agents can create complex dynamics. Agents using independent reward updates frequently converged to Pareto-dominated equilibria. Successful approaches employed:
with inter-agent reward sensitivity β. This succeeded when:
- Agents had aligned initial reward functions
- The adaptation rate β decayed appropriately
- Communication channels provided sufficient signaling
Limitations in Sparse-Reward Environments
Montezuma's Revenge benchmarks revealed fundamental challenges when evolving rewards from extremely sparse signals. Techniques like:
using state embeddings ϕ often failed because:
- Novelty rewards didn't correlate with task completion
- The embedding space captured irrelevant features
- Adaptation caused catastrophic forgetting of critical skills
Recommendations for Practical Implementation
Empirical studies suggest these principles for evolving reward functions:
- Conservation of reward mass: Maintain ∫Rt(s)ds ≈ C across adaptations
- Slow adaptation rates: Keep ΔR/Δt below the policy's learning rate
- Regularization: Penalize large reward function changes ‖Rt+1-Rt‖
4. Frameworks for Implementing Dynamic Reward Functions
4.1 Frameworks for Implementing Dynamic Reward Functions
Dynamic reward functions adapt to an agent's learning progress, environmental changes, or shifting objectives. Unlike static reward functions, they require frameworks that support real-time updates, parameter tuning, and feedback integration. Three primary approaches dominate this space: meta-learning-based adaptation, human-in-the-loop tuning, and automated reward shaping via intrinsic motivation.
Meta-Learning-Based Adaptation
Meta-learning frameworks treat the reward function as a learnable component, optimizing it alongside the policy. The reward function \( R_{\phi}(s, a) \) is parameterized by \(\phi\), which is updated via gradient descent to maximize a meta-objective, such as task performance or sample efficiency. The meta-update rule is:
Here, \(\alpha\) is the meta-learning rate, and \(\pi_{\theta}\) is the policy being trained. Practical implementations often use bi-level optimization, where the inner loop trains the policy, and the outer loop adjusts \(\phi\). For example, in robotics, this enables reward functions to evolve as the agent masters subtasks like grasping or navigation.
Human-in-the-Loop Tuning
Interactive frameworks allow human operators to refine rewards based on observed behavior. Techniques like preference-based learning (e.g., Deep Reinforcement Learning from Human Preferences) use pairwise comparisons to iteratively adjust \(R(s, a)\). The reward update follows:
where \(\beta\) scales human feedback. Real-world applications include autonomous driving, where safety-critical rewards are adjusted based on driver interventions.
Automated Reward Shaping
Intrinsic motivation mechanisms autonomously modify rewards using curiosity or novelty signals. A common formulation combines extrinsic rewards \(R_{ext}\) with intrinsic rewards \(R_{int}\):
The weighting factor \(\lambda(t)\) decays over time to phase out exploration bias. Variants include:
- Prediction error-based: \(R_{int} = \|f(s_{t+1}) - \hat{f}(s_{t+1})\|^2\), where \(f\) is a dynamics model.
- Count-based: \(R_{int} = 1/\sqrt{N(s)}\), with \(N(s)\) tracking state visits.
In OpenAI's Hide-and-Seek, intrinsic rewards for tool discovery led to emergent strategic behaviors. The framework dynamically reduced \(\lambda(t)\) as agents mastered subgoals.

4.2 Debugging and Evaluating Evolving Reward Systems
Evolving reward functions introduce unique challenges in reinforcement learning (RL) due to their dynamic nature. Traditional debugging techniques for static reward functions often fail when the reward landscape shifts during training. To systematically diagnose issues, we must analyze three key components: reward drift, policy adaptation lag, and credit assignment consistency.
Detecting Reward Drift
Reward drift occurs when the agent's policy exploits the evolving reward function in unintended ways, leading to degenerate solutions. To quantify drift, measure the KL-divergence between the expected and observed reward distributions:
where P is the designer's intended reward distribution and Q is the empirical reward distribution from agent trajectories. A divergence threshold (typically 0.2-0.5 nats) signals problematic drift.
Policy Adaptation Metrics
When rewards evolve faster than the policy can adapt, the agent exhibits suboptimal behavior. Track the adaptation gap:
where k represents the policy update interval. For stable learning, Δt should converge to zero as training progresses.
Credit Assignment Analysis
Evolving rewards complicate temporal credit assignment. Use counterfactual advantage estimation to verify if the agent correctly attributes rewards to actions. Compute the advantage function Aπ(s,a) using generalized advantage estimation (GAE):
Visualize the advantage matrix across state-action pairs to identify misattributions. Persistent negative advantages for optimal actions indicate reward function issues.
Diagnostic Tools
Implement these practical debugging techniques:
- Reward Landscape Visualization: Plot iso-reward contours for fixed policies across training iterations
- Policy Response Heatmaps: Track how policy logits change in response to reward updates
- Trajectory Clustering: Group episodes by reward structure to detect unintended behavioral modes
Evaluation Protocol
For rigorous assessment of evolving reward systems:
- Maintain a fixed reference policy trained on initial rewards as a baseline
- Compute the normalized policy improvement metric:
where J(π) is the expected return and π* is the optimal policy. Values below 0 indicate catastrophic forgetting.
Real-World Case Study
In robotic control with curriculum learning, reward functions often evolve from sparse to dense formulations. A common failure mode occurs when the agent overfits to early reward stages. The solution involves:
- Adding reward shaping invariance tests that verify policy consistency under potential-based shaping
- Implementing policy distillation between reward updates to preserve learned skills
- Using adaptive reward clipping to maintain stable gradient magnitudes

4.3 Best Practices for Scalability and Robustness
Modular Reward Decomposition
Scalable reward functions must decompose complex objectives into modular sub-rewards, enabling incremental learning and credit assignment. Given a global reward R, decompose it into k sub-rewards ri with learnable weights wi:
Weights wi can be adapted via gradient-based meta-learning:
This approach prevents reward hacking by isolating contributions from distinct behavioral components.
Curriculum Learning Integration
Robustness requires phased difficulty progression. Implement a curriculum scheduler C(τ) that dynamically adjusts reward thresholds based on agent performance:
where Pt is the success rate over a sliding window, δ is the target performance, and α, β control adaptation speed. This prevents plateaus in sparse-reward environments.
Non-Stationarity Mitigation
For environments with shifting dynamics, employ predictive reward normalization. Maintain a running estimate of reward statistics:
Normalized rewards r̃t = (rt - μ̂t)/σ̂t maintain consistent gradient magnitudes across training phases.
Multi-Objective Pareto Optimization
When conflicting sub-rewards exist, model them as a Pareto front. The reward vector R = [r1,..., rk] induces a partial ordering over policies. Use constrained policy optimization:
where εi are minimum performance thresholds. This prevents catastrophic neglect of secondary objectives.
Adversarial Reward Robustness
To defend against reward function exploitation, train with adversarial perturbations δ bounded by Lp-norm constraints:
Regularize the policy using worst-case rewards to improve generalization. This is particularly critical for real-world deployment where reward misspecification is common.
Distributed Reward Learning
For large-scale systems, implement distributed reward updates using parameter servers. Each worker j computes local gradients g(j)t, which are aggregated asynchronously:
Use importance weighting to correct for policy divergence across workers. This architecture enables training on millions of diverse environment instances while maintaining reward consistency.

5. Bias and Fairness in Evolving Reward Functions
5.1 Bias and Fairness in Evolving Reward Functions
Evolving reward functions introduce unique challenges in maintaining fairness and mitigating bias, particularly as the agent's learning process dynamically alters the reward landscape. Unlike static reward functions, where bias can be analyzed and corrected upfront, evolving rewards require continuous monitoring and adaptation to prevent unintended discriminatory outcomes.
Sources of Bias in Dynamic Reward Systems
Bias in evolving reward functions can emerge from multiple sources:
- Initial Reward Shaping: If the initial reward function encodes implicit biases (e.g., favoring certain demographic groups in a recommendation system), these biases can propagate and amplify as the reward function evolves.
- Feedback Loops: The agent's actions influence the environment, which in turn affects future rewards. Biased actions can create self-reinforcing loops, exacerbating disparities over time.
- Data Distribution Shifts: As the agent explores, the distribution of states and actions changes, potentially exposing underrepresented regions where the reward function behaves unfairly.
Mathematically, bias can be formalized as a discrepancy in expected rewards across different subpopulations. Let S be the state space partitioned into subgroups S1, S2, ..., Sn. The bias B for a reward function R is:
Fairness Constraints in Adaptive Rewards
To enforce fairness, constraints can be integrated into the reward adaptation mechanism. One approach is to formulate a constrained optimization problem where the reward function maximizes expected return while minimizing bias:
Here, π* is the optimal policy under R, and ε is a fairness tolerance threshold. Lagrangian relaxation can be used to convert this into an unconstrained problem:
Dynamic Fairness-Aware Reward Adaptation
Practical implementations often use gradient-based methods to adapt rewards while monitoring fairness metrics. For a parameterized reward function Rθ, the update rule becomes:
where J(θ) is the expected return and α is the learning rate. This requires efficient estimation of the bias gradient, which can be approximated using sampled trajectories from different subpopulations.
Case Study: Bias Mitigation in Recommender Systems
A real-world example involves recommendation algorithms where evolving rewards optimize for user engagement. Without fairness constraints, these systems often amplify popularity bias, favoring already dominant content. By dynamically adjusting rewards to promote diversity (e.g., incorporating entropy regularization over recommended items), the system can maintain engagement while reducing bias:
Here, η controls the strength of diversity promotion, and P(a|s) is the recommendation probability distribution.
5.2 Long-Term Impacts on Agent Behavior
The evolution of reward functions during agent learning introduces complex dynamics that shape long-term behavior. Unlike static reward functions, adaptive rewards create a feedback loop where the agent's policy influences future reward structures, which in turn guide further policy updates. This recursive relationship can lead to emergent phenomena such as reward hacking, distributional shift, or unintended convergence to suboptimal equilibria.
Mathematical Formulation of Evolving Rewards
Consider a Markov Decision Process (MDP) with a state space S, action space A, and transition dynamics P(s'|s,a). The traditional reward function R(s,a) is replaced by a parameterized family R_θ(s,a), where θ evolves based on the agent's learning trajectory. The joint optimization can be expressed as:
subject to constraints ensuring θ remains within a feasible set Θ that prevents degenerate solutions. The gradient update for θ often incorporates a meta-objective, such as:
where Qπ(s,a) is the state-action value function under policy π. This formulation reveals how reward updates depend on the agent's current value estimates, creating a bidirectional coupling between policy and reward learning.
Behavioral Consequences of Reward Adaptation
Three primary long-term effects emerge from evolving reward functions:
- Non-Markovian Dependencies: As rewards adapt to the agent's behavior, the environment effectively becomes non-Markovian from the agent's perspective, since historical policy choices influence current rewards.
- Path-Dependent Convergence: The system exhibits hysteresis, where final behavior depends strongly on the initial conditions and exploration trajectory.
- Reward Distortion: The agent may discover policies that artificially inflate rewards without achieving true objectives, a manifestation of Goodhart's Law in RL.
Case Study: Reward Function Drift in Continual Learning
In a continual learning scenario where an agent faces sequentially changing tasks, an evolving reward function can either mitigate catastrophic forgetting or exacerbate it. Consider a neural network policy where the reward function is represented as:
Here, fφ provides immediate task rewards while gψ computes a history-dependent bonus based on recent trajectories τ1:k. Experiments show that careful tuning of the adaptation rate λ is critical—too rapid adaptation leads to reward instability, while too slow adaptation fails to prevent forgetting.
Empirical Observations from Multi-Agent Systems
In multi-agent environments, evolving rewards create additional complexity. When two agents A and B learn with interdependent reward functions RθA(s,aA,aB) and RθB(s,aB,aA), the system dynamics resemble a continuous game where each player's strategy affects the other's payoff structure. Theoretical analysis reveals:
The cross-term with coefficient ε captures how one agent's reward parameters influence the other's learning, leading to either cooperative adaptation or destructive interference depending on the sign and magnitude of ε.
Mitigation Strategies for Unintended Consequences
Several approaches have proven effective in managing long-term behavioral impacts:
- Regularization Constraints: Adding a KL-divergence penalty DKL(Rθ || Rθ0) prevents excessive reward distortion.
- Two-Timescale Learning: Updating θ at a slower rate than the policy parameters stabilizes the coupled dynamics.
- Counterfactual Reward Modeling: Maintaining a baseline reward function enables detection of harmful adaptations.

5.3 Open Research Questions and Emerging Trends
Non-Stationary Reward Learning
A fundamental challenge in adaptive reward functions is the non-stationarity introduced when rewards evolve alongside the agent's policy. Traditional reinforcement learning assumes a fixed reward function R(s, a), but in dynamic settings, the optimal reward function R*(s, a) may shift as the agent's capabilities improve. This creates a coupled optimization problem:
where θ parameterizes the reward function and π* is the optimal policy under Rθ. Recent work in inverse reinforcement learning (IRL) has explored gradient-based methods to update θ concurrently with policy optimization, but stability remains an open issue.
Multi-Objective Reward Adaptation
Emerging techniques formulate reward adaptation as a Pareto optimization problem, where the reward function must balance competing objectives (e.g., exploration vs. exploitation, short-term vs. long-term gains). The dynamic weight adjustment can be modeled as:
with weights wi(π) conditioned on the agent's current policy. Research frontiers include meta-learning approaches to predict optimal weight trajectories and game-theoretic formulations where the reward function acts as an adversarial player.
Interpretability vs. Performance Trade-offs
As reward functions become increasingly parameterized (e.g., via deep neural networks), interpretability declines—a critical concern in safety-sensitive domains. Current approaches attempt to:
- Constrain reward networks to sparse linear combinations of human-understandable features
- Employ symbolic regression to extract interpretable reward formulas from neural networks
- Develop explanation frameworks that map complex rewards to human-aligned concepts
However, no method yet achieves both the expressivity of deep reward functions and the interpretability of hand-designed rewards.
Emerging Architectures
Three novel architectures show promise for dynamic reward learning:
- Hypernetwork-based rewards: A meta-network generates the reward function's weights based on the agent's learning stage
- Memory-augmented rewards: External memory stores past reward adjustments, enabling long-term credit assignment
- Physics-inspired rewards: Potential-based shaping functions derived from conservation laws in dynamical systems
Frontier Challenges
Key unsolved problems include:
- Catastrophic forgetting: Reward updates may erase previously learned behaviors
- Distributional shift: The agent's changing policy induces non-IID reward samples
- Multi-agent alignment: Coordinating reward adaptation across interacting agents
Recent work in continual learning and off-policy correction offers partial solutions, but fundamental limitations persist in non-Markovian settings.

6. Key Research Papers and Seminal Works
6.1 Key Research Papers and Seminal Works
- A review of research on reinforcement learning algorithms for multi ... — MARL combines Collaborative Learning (CL) and RL to address the collaboration and competition problem in MAS, whereby each agent interacts with other agents by sensing the environment to learn the best behavioral strategies to optimize overall performance.CL focuses on cooperation and coordination between multi-agent, while RL explores how agents can derive rewards from the environment and ...
- Inducing structure in reward learning by learning features — Whether it is semi-autonomous driving (Sadigh et al., 2016), recommender systems (Ziebart et al., 2008), or household robots working in close proximity with people (Jain et al., 2015), reward learning can greatly benefit autonomous agents to generate behaviors that adapt to new situations or human preferences.Under this framework, the robot uses the person's input to learn a reward function ...
- Learning reward functions from diverse sources of human feedback ... — Prior work has extensively studied learning reward functions using a single source of information, e.g., ordinal data (Chu and Ghahramani, 2005) or human corrections (Bajcsy et al., 2018, 2017; Li et al., 2021b). Other works attempted to incorporate expert assessments of trajectories (Shah and Shah, 2020). More related to our work, we focus on ...
- PDF Potential-Based Reward Shaping for Knowledge-Based Multi-Agent ... — number of suboptimal behaviours tried and, therefore, speed up the rate of learning. Potential-based reward shaping is a method of providing this knowledge to an agent by additional rewards. Furthermore, if the agent is alone in the environment, it is guaranteed to learn the same behaviour both with and without potential-based reward shaping.
- Designing Reward Functions Using Active Preference Learning for ... — This study presents a method based on active preference learning to overcome the challenges of designing reward functions for autonomous navigation. Results obtained from training with artificially designed reward functions may not accurately reflect human intentions. We focus on the limitations of traditional reward functions, which often fail to facilitate complex tasks in continuous state ...
- PDF Reward Redistribution Mechanisms in Multi-agent Reinforcement Learning — In typical Multi-Agent Reinforcement Learning (MARL) settings, each agent acts to maximize its individual reward objective. How-ever, for collective social welfare maximization, some agents may need to act non-selfishly. We propose a reward shaping mechanism using extrinsic motivation for achieving modularity and increased
- PDF Quantifying Differences in Reward Functions - University of California ... — ulation tasks [15]. Unfortunately, the reward functions for most real-world tasks are difficult or impossible to procedurally specify. Even a task as simple as peg insertion from pixels has a non-trivial reward function that must usually be learned [21, IV.A]. Most real-world tasks have far more complex reward functions than this.
- Reward Machines: Exploiting Reward Function Structure in Reinforcement ... — outputs the reward function the agent should use at that time. For example, we might construct a reward machine for \delivering co ee to an o ce" using two states. In the rst state, the agent does not receive any rewards, but it moves to the second state whenever it gets the co ee. In the second state, the agent gets rewards after delivering ...
- Reward criteria impact on the performance of reinforcement learning ... — In reinforcement learning, an agent takes action at every time step (follows a policy) in an environment to maximize the expected cumulative reward. Therefore, the shaping of a reward function plays a crucial role in an agent's learning. Designing an optimal reward function is not a trivial task.
- Accelerating Reinforcement Learning using EEG-based implicit human ... — Before RL agent starts learning, we ask the human subjects to observe a number of trajectories, and record their implicit feedbacks in the form of ErrP on corresponding state-action pair in a dataset. Then we learn the auxiliary reward function r a (·, ·) from these trajectories labeled by human feedback. During the RL training, the learned ...
6.2 Recommended Books and Online Resources
- PDF MULTI-AGENT RE INFORCEMENT LEAR NING - marl-book — 6.2 Temporal-Difference Learning for Games: Joint-Action Learning118 6.2.1 Minimax Q-Learning121 6.2.2 Nash Q-Learning123 6.2.3 Correlated Q-Learning124 6.2.4 Limitations of Joint-Action Learning125 6.3 Agent Modeling127 6.3.1 Fictitious Play128 6.3.2 Joint-Action Learning with Agent Modeling131 6.3.3 Bayesian Learning and Value of Information134
- PDF The Optimal Reward Problem: Designing E ective Reward for Bounded Agents — agent, while a separate agent reward function is used to guide agent behavior. The designer can then solve the Optimal Reward Problem (ORP): choose the agent reward function which leads to the greatest expected reward for the designer. The second contribution is the demonstration through examples that good reward functions are chosen by ...
- MuDE: Multi-agent decomposed reward-based exploration — Multi-agent reinforcement learning (MARL) has demonstrated its efficiency in a wide range of challenging applications, such as humanoid robots, autonomous cars, and sensor networks (Dinneweth et al., 2022, Hüttenrauch et al., 2017, Ye et al., 2015).In such applications often modeled as a cooperative MARL, agents jointly optimize a centralized value function based on the rewards given for ...
- Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion ... — to optimize downstream reward functions. While diffusion models are widely known to provide excellent generative modeling capability, practical applications in domains such as biology ... (2023);Uehara et al.(2024)), aiming to optimize downstream reward functions. RL is a machine learning paradigm where agents learn to make sequential decisions ...
- Designing Reward Functions Using Active Preference Learning for ... — This study presents a method based on active preference learning to overcome the challenges of designing reward functions for autonomous navigation. Results obtained from training with artificially designed reward functions may not accurately reflect human intentions. We focus on the limitations of traditional reward functions, which often fail to facilitate complex tasks in continuous state ...
- Online learning of shaping rewards in reinforcement learning — Reinforcement learning (RL) is a popular method to design autonomous agents that learn from interactions with the environment, or, in more mathematical terms, from repeated simulation (Bertsekas, 2007).In contrast to supervised learning (Mitchell, 1997), RL methods do not rely on instructive feedback; i.e., the agent is not informed as to what the best action is in a given situation.
- Tiered Reward: Designing Rewards for Specification and Fast Learning of ... — They reveal the structure of the reward function to the RL agent to support decomposition of complex tasks. Our focus on how to provide incentives for specific outcomes is complementary and the two approaches can be used in concert. ... Policy teaching through reward function learning. In Proceedings of the 10th ACM conference on Electronic ...
- PDF Quantifying Differences in Reward Functions - University of California ... — ulation tasks [15]. Unfortunately, the reward functions for most real-world tasks are difficult or impossible to procedurally specify. Even a task as simple as peg insertion from pixels has a non-trivial reward function that must usually be learned [21, IV.A]. Most real-world tasks have far more complex reward functions than this.
- STRIDE: Automating Reward Design, Deep Reinforcement Learning Training ... — These reward functions are evaluated by training agents in a simulated environment, with their performance scored using a predefined fitness function. STRIDE iteratively refines the reward functions by incorporating textual feedback generated from training results. This closed-loop process continues until an optimal reward function is obtained.
- Online Intrinsic Rewards for Decision Making Agents from Large Language ... — For online reward learning via LLM annotations in this system 1 1 1 The prior work by Klissarov et al. in the offline setting was able to connect their intrinsic reward function into Sample Factory's APPO implementation with a few lines of code, as it just needed to load the PyTorch module of the learned reward function., we added 1) an LLM ...
6.3 Communities and Conferences for Ongoing Learning
- Reinforcement learning - Wikipedia — Reinforcement learning (RL) is an interdisciplinary area of machine learning and optimal control concerned with how an intelligent agent should take actions in a dynamic environment in order to maximize a reward signal. Reinforcement learning is one of the three basic machine learning paradigms, alongside supervised learning and unsupervised learning. ...
- Multi-Agent Deep Reinforcement Learning for Blockchain-Based ... - MDPI — To refine the valuation function, the Multi-agent Proximal Policy Optimization (MAPPO) , an algorithm based on reinforcement learning, facilitates the training of numerous agents within settings where rewards serve as feedback. Within the POMDP framework, each agent assesses the current state of the system and strategizes to enhance rewards ...
- Designing Reward Functions Using Active Preference Learning for ... — This study presents a method based on active preference learning to overcome the challenges of designing reward functions for autonomous navigation. Results obtained from training with artificially designed reward functions may not accurately reflect human intentions. We focus on the limitations of traditional reward functions, which often fail to facilitate complex tasks in continuous state ...
- Bottom-up multi-agent reinforcement learning by reward shaping for ... — A multi-agent system (MAS) is expected to be applied to various real-world problems where a single agent cannot accomplish given tasks. Due to the inherent complexity in the real-world MAS, however, manual design of group behaviors of agents is intractable. Multi-agent reinforcement learning (MARL), which is a framework for multiple agents in the same environment to learn their policies ...
- Active reward learning with a novel acquisition function — Reward functions are an essential component of many robot learning methods. Defining such functions, however, remains hard in many practical applications. For tasks such as grasping, there are no reliable success measures available. Defining reward functions by hand requires extensive task knowledge and often leads to undesired emergent behavior. We introduce a framework, wherein the robot ...
- (PDF) The theory of social functions: challenges for computational ... — Restating Let us now look at the same phenomenon with another perspective able to enlightening another — concurrent — mechanism.20 Even without postulating any reinforcement and learning by the agent, an effect that maintains or re-creates those contextual conditions that lead to that action, maintains or increases the probability for its ...
- Adaptive multi-agent reinforcement learning for flexible resource ... — An effective building coordination scheme can be established by carefully crafting the reward function and implementing a sound cooperative mechanism. Moreover, the learning process of the algorithm, enriched by extensive interactions with real-world environmental data, enables the capturing and understanding of system uncertainties.
- Multi-objective reinforcement learning for designing ethical multi ... — This paper tackles the open problem of value alignment in multi-agent systems. In particular, we propose an approach to build an ethical environment that guarantees that agents in the system learn a joint ethically-aligned behaviour while pursuing their respective individual objectives. Our contributions are founded in the framework of Multi-Objective Multi-Agent Reinforcement Learning ...
- Learning reward functions from diverse sources of human feedback ... — Reward functions are a common way to specify the objective of a robot. As designing reward functions can be extremely challenging, a more promising approach is to directly learn reward functions ...
- Frontiers | Decentralized multi-agent reinforcement learning based on ... — which computes a deterministic function f Π (s, ζ) that depends on the state s, policy parameters Π, and independent noise vector ζ drawn from a fixed distribution, e.g., mean free Gaussian noise. In contrast to DDPG, this parameterized policy is also squashed via a tanh function to the bounds of the action space, thus resulting in valid samples that can be used to generate a stochastic ...








