This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| SARSA | |
|---|---|
| Name | SARSA |
| Caption | On-policy temporal-difference learning |
| Developer | Richard S. Sutton and Andrew G. Barto |
| Introduced | 1990s |
| Paradigm | Reinforcement learning |
| Input | State, action, reward, next state, next action |
| Output | Action-value function (Q) |
SARSA. SARSA is an on-policy temporal-difference control algorithm in reinforcement learning devised to estimate action-value functions and learn policies through interaction with an environment. It combines ideas from dynamic programming, Monte Carlo methods, and temporal-difference learning to update estimates from sampled transitions and is closely associated with foundational work by researchers at institutions such as University of Alberta, University of Massachusetts Amherst, and laboratories including MIT Computer Science and Artificial Intelligence Laboratory and DeepMind. SARSA has been applied across robotics, games, and control tasks studied at venues like the International Joint Conference on Artificial Intelligence, NeurIPS, ICML, and AAAI.
SARSA originated from the reinforcement learning literature alongside algorithms such as Q-learning, Temporal difference learning, and algorithms influenced by the textbook by Sutton and Barto and workshops at NIPS 1987 and ICML 1990. It operates within the formalism of Markov decision processs and uses sample transitions akin to methods employed by researchers at Bell Labs and projects at AT&T Bell Laboratories and research groups at Stanford University. Early comparisons involved experiments replicating results from groups at Carnegie Mellon University, University of California, Berkeley, and collaborations with researchers at IBM Research. SARSA's on-policy nature contrasts with off-policy methods evaluated in studies at DeepMind and theoretical analyses at institutions like Princeton University.
SARSA updates an action-value estimate Q(s,a) using a quintuple (s, a, r, s', a') sampled during interaction; the update rule resembles procedures used in algorithms researched at Bellcore and in engineering groups at Toyota Research Institute. The iterative update is analogous to bootstrapping schemes discussed in literature from Cambridge University and formalized in conferences at COLT and UAI. Implementations of SARSA have been benchmarked in environments devised by labs such as OpenAI and evaluated against baselines from teams at Google DeepMind, Microsoft Research, and universities including Yale University and University of Oxford.
Convergence proofs for SARSA under conditions like decaying learning rates, exploration policies such as ε-greedy, and finite MDPs were developed building on stochastic approximation theory from researchers at MIT, Harvard University, and work by authors affiliated with INRIA and ETH Zurich. Theoretical analysis often references results from the stochastic optimization literature at Courant Institute and probabilistic tools used in studies at Columbia University and University of Chicago. Comparisons with convergence properties of Q-learning and policy iteration draw on seminal texts and papers associated with Royal Society-supported projects and grants from agencies such as NSF and DARPA.
Many extensions modify SARSA with eligibility traces (SARSA(λ)), function approximation using linear or nonlinear approximators inspired by work at Google Brain and Facebook AI Research, and adaptations for continuous action spaces informed by research at Caltech and UC San Diego. Actor-critic hybrids, prioritized sweeping, and expected SARSA variants were proposed in the literature developed in groups at University of Toronto, University of Waterloo, and labs like Adobe Research. Recent deep reinforcement learning adaptations replace tabular Q with neural networks as explored by teams at DeepMind, OpenAI, Berkeley AI Research, and Carnegie Mellon University.
SARSA has been applied to control problems in robotics labs such as KUKA Research, Honda Research Institute, and projects at NASA and European Space Agency for autonomous systems. It has been used in game-playing agents inspired by experiments at University of Alberta and competitions like the RoboCup and has informed traffic signal control studies led by groups at Toyota Technological Institute, Imperial College London, and city projects with Siemens. SARSA-based controllers appear in industrial automation at companies like Siemens and Bosch and in energy management research at ENERGY STAR-affiliated projects and academic collaborations with ETH Zurich.
Practical implementations of SARSA require choices about exploration schedules, learning-rate decay, state representation, and feature construction—topics addressed in tutorials at NeurIPS, ICML, and summer schools at CIFAR and Mathematical Sciences Research Institute. Software libraries and frameworks that include SARSA examples are maintained by organizations such as OpenAI, TensorFlow teams at Google, and contributors from GitHub repositories associated with researchers at University of Illinois Urbana-Champaign and Purdue University. Empirical best practices often follow benchmarking suites developed by groups at Stanford University, UC Berkeley, and consortiums supported by ERC grants and industrial partners like Intel and NVIDIA.
Category:Reinforcement learning algorithms