LLMpediaThe first transparent, open encyclopedia generated by LLMs

Sequence-to-sequence learning

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: Google Research New York Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

Sequence-to-sequence learning
NameSequence-to-sequence learning
AltSeq2seq
Introduced2014
ContributorsGeoffrey Hinton, Yoshua Bengio, Yann LeCun, Ilya Sutskever, Oriol Vinyals, Kyunghyun Cho
ParadigmSupervised learning
DomainsGoogle, Microsoft, Facebook, OpenAI, DeepMind

Sequence-to-sequence learning Sequence-to-sequence learning is a class of machine learning methods that maps variable-length input sequences to variable-length output sequences, widely used in tasks such as machine translation, speech recognition, and text summarization. Originating from advances associated with researchers like Ilya Sutskever and organizations such as Google and OpenAI, the paradigm unified architectures employed by teams at Microsoft Research, Facebook AI Research, and DeepMind. Sequence-to-sequence approaches frequently leverage innovations from work at University of Montreal, Massachusetts Institute of Technology, Courant Institute, Stanford University, and industrial labs including IBM Research.

Introduction

Early demonstrations of sequence-to-sequence capabilities built upon breakthroughs by Geoffrey Hinton, Yoshua Bengio, and Yann LeCun in representation learning, and were popularized by papers from Ilya Sutskever, Oriol Vinyals, and Kyunghyun Cho. Pivotal implementations integrated recurrent architectures championed at University of Toronto and attention mechanisms influenced by research groups at Google Brain and Google DeepMind. The approach bridged contributions from applied groups at Amazon Web Services, Apple Inc., NVIDIA, and academic labs at Princeton University, University of Oxford, University of Cambridge, and Carnegie Mellon University.

Model Architectures

Standard architectures include encoder–decoder models derived from recurrent neural networks promoted by researchers at NYU, ETH Zurich, and University College London. Variants incorporate long short-term memory units associated with work by Sepp Hochreiter and Jürgen Schmidhuber, gated recurrent units introduced in collaborations at Kyoto University and Samsung Research, and convolutional sequence models inspired by research teams at Facebook AI Research and Google. The transformer architecture, introduced by authors affiliated with Google Brain, replaced recurrence in many systems and influenced deployments at OpenAI, Microsoft Research, Alibaba Group, Tencent, and Baidu Research. Hybrid systems combine ideas from MIT-IBM Watson AI Lab, Allen Institute for AI, and Salesforce Research.

Training and Optimization

Training methods draw on stochastic gradient descent used in frameworks developed by TensorFlow teams at Google, PyTorch contributions from Facebook, and compilation efforts at Intel Corporation and AMD. Techniques such as teacher forcing were analyzed in publications from University of Edinburgh and University of Washington, while curriculum learning concepts traced to work at UC Berkeley. Regularization and optimization strategies referenced in studies at Stanford University, Harvard University, University of California, Berkeley, and Caltech include label smoothing, dropout, Adam optimizer, and learning rate scheduling pioneered in labs at Google AI, OpenAI, and DeepMind.

Applications

Sequence-to-sequence models have been applied to machine translation products from Google Translate, Microsoft Translator, and initiatives at Baidu Translate, to automatic speech recognition systems developed at Nuance Communications and Amazon, and to dialogue agents showcased by OpenAI ChatGPT, Google Assistant, and Apple Siri. Other deployments include code generation efforts at GitHub Copilot in collaboration with OpenAI, summarization engines from DeepMind, biomedical information extraction projects at National Institutes of Health, and captioning systems implemented by teams at YouTube and Facebook.

Evaluation Metrics and Benchmarks

Benchmarking has relied on datasets and leaderboards curated by entities such as WMT and groups at Stanford Question Answering Dataset and evaluations framed by initiatives at BLEU creators, ROUGE contributors, and metrics informed by research at GLUE and SuperGLUE teams. Empirical comparisons commonly use corpora maintained by LDC and testbeds provided by Common Crawl, while academic competitions at NeurIPS, ICML, ACL, and EMNLP have driven metric development. External audits and reproducibility studies have emerged from collaborations among Carnegie Mellon University, ETH Zurich, and University of Pennsylvania.

Challenges and Limitations

Sequence-to-sequence systems face challenges identified by researchers at MIT, Stanford University, and Princeton University, including exposure bias, catastrophic forgetting observed in studies from Google Research and Facebook AI Research, and robustness issues examined by teams at Microsoft Research and OpenAI. Data scarcity and domain adaptation remain concerns highlighted by projects at DARPA and European Commission initiatives, while computational cost and environmental impact have been critiqued by researchers at University of Massachusetts Amherst and University of Copenhagen. Ethical and safety considerations have been raised in reports involving UNESCO, OECD, AI Now Institute, and civil society partners.

Extensions and Variants

Extensions include unsupervised and semi-supervised frameworks advanced at DeepMind and Facebook AI Research, multimodal seq2seq models pursued at Google Research and OpenAI, and reinforcement learning integrations explored by teams at DeepMind and DeepMind x Google Deep Learning collaborations. Other variants comprise pointer-generator networks from University of California, Berkeley, variational encoder–decoder models studied at University of Toronto, and sparse attention mechanisms developed by researchers at Stanford University and Carnegie Mellon University. Transfer learning and pretraining paradigms used by OpenAI, Google, Facebook, and Microsoft continue to extend sequence-to-sequence capabilities.

Category:Machine learning