LLMpediaThe first transparent, open encyclopedia generated by LLMs

multilayer perceptron

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: Parallel Distributed Processing Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

multilayer perceptron
NameMultilayer perceptron
InventorsFrank Rosenblatt
Introduced1960s
TypeArtificial neural network
ApplicationsPattern recognition, function approximation, classification

multilayer perceptron

Introduction

A multilayer perceptron is a class of feedforward artificial neural network architecture developed in the context of early computing by Frank Rosenblatt and elaborated through work at Cornell University, Stanford University, and Bell Labs during the mid-20th century; it influenced research at MIT, Carnegie Mellon University, and IBM before resurgence via breakthroughs at University of Toronto, University College London, and Google DeepMind. The model gained prominence following theoretical results by researchers associated with University of California, Berkeley, University of Edinburgh, and University of Montreal, and was central to experiments at laboratories such as Bell Labs, Xerox PARC, and Microsoft Research.

Architecture

A multilayer perceptron comprises an input layer, one or more hidden layers, and an output layer, and its topology was studied in collaboration between teams at Bell Labs, AT&T Laboratories, and Hewlett-Packard; practical implementations appeared in systems developed at Sun Microsystems, Intel, and NVIDIA. The units in each layer perform weighted summations with biases and pass results through nonlinear functions, an idea refined by researchers at Princeton University, Harvard University, and Columbia University and implemented on hardware platforms from IBM Watson projects and Cray Research installations. Architectures vary from shallow configurations used in studies at Los Alamos National Laboratory to deep stacks explored by groups at Facebook AI Research, DeepMind, and OpenAI.

Training and Learning Algorithms

Training is typically performed by supervised learning using backpropagation combined with gradient-based optimizers like stochastic gradient descent popularized in research at Stanford University, NYU, and ETH Zurich; extensions include momentum methods influenced by work at University of Toronto and adaptive schemes such as Adam developed by researchers affiliated with Google, University of California, Berkeley, and Microsoft Research. Curriculum learning schedules were explored in projects at Google Brain and DeepMind, while regularization strategies were compared across experiments at University of Oxford and Imperial College London. Early convergence analyses drew on mathematical techniques from Princeton University and Yale University; large-scale training pipelines were engineered by teams at Amazon Web Services and Alibaba.

Activation Functions and Regularization

Common activation functions include sigmoidal functions formalized in discussions at Bell Labs and hyperbolic tangent variants used in implementations at Xerox PARC and Hewlett-Packard, while rectified linear units popularized through work at University of Toronto, University of Montreal, and Google altered practice in projects at Facebook, Microsoft, and NVIDIA. Regularization techniques such as weight decay, dropout introduced by researchers at Georgetown University and University College London, and batch normalization from teams at University of Toronto and Stanford University were validated in benchmarks run by OpenAI, DeepMind, and Facebook AI Research. Architectural choices were influenced by empirical studies from Carnegie Mellon University and theoretical analyses from Columbia University and University of Illinois Urbana-Champaign.

Variants and Extensions

Variants include deep multilayer perceptrons explored at University of Toronto, residual connections inspired by work at Microsoft Research Asia and Chinese University of Hong Kong, and sparse or Bayesian extensions linked to research at University of Cambridge and University of Washington. Integrations with convolutional components were advanced at NYU and University of Oxford for image tasks studied at Stanford University and Caltech, while recurrent hybrids were investigated at University of Montreal and MILA. Meta-learning and transfer learning adaptations were developed in collaborations involving DeepMind, OpenAI, and Google Brain.

Applications

Multilayer perceptrons have been applied to classification and regression problems in systems deployed by IBM, Microsoft, and Amazon; they were used in speech recognition research at Bell Labs, machine translation studies at Google Translate teams connected to Google Research, and diagnostic tools developed at Mayo Clinic and Johns Hopkins University. Financial modelling experiments were carried out in collaborations with Goldman Sachs and JPMorgan Chase, while scientific applications involved simulations at CERN, climate modelling projects with NASA, and genomics pipelines at Broad Institute. Industrial deployments were reported by GE, Siemens, and Bosch.

Theoretical Properties and Limitations

Universal approximation results attributed to researchers at University of Illinois Urbana-Champaign and Ohio State University established that sufficiently large multilayer networks can approximate measurable functions, a theorem discussed in seminars at Princeton University and Harvard University; however, complexity bounds and generalization behavior were further analyzed in studies from ETH Zurich, University of Cambridge, and MIT. Limitations include susceptibility to adversarial examples explored by teams at Google Brain and OpenAI, optimization difficulties studied at Stanford University and Columbia University, and scalability constraints addressed by hardware efforts at NVIDIA and Intel Research.

Category:Artificial neural networks