My research has moved between computational neuroscience, biologically plausible learning algorithms, and large-scale foundation models. Broadly I have been attempting to understand how intelligence works in both brains and machines, and specifically understanding questions such as how learning and inference can be implemented in local and asynchronous biological systems, how associative memories work, the underlying origins of exploration in RL, and most recently how architectures, data, and training systems come together to build highly capable foundation models across a wide range of modalities.
Google Scholar · GitHub · Full publication list
Selected research
The sections below collect a small number of papers into the main research programs I have worked on. They are intended as an entry point; the complete reverse-chronological publication list is below.
Foundation models: architectures, data, and scaling
Since cofounding Zyphra, my work has increasingly focused on building and leading research across foundation-model architectures, data, training systems, as well as trying to attack fundamental questions of continual learning and long term memory. A recurring question is which constraints actually limit scaling, and how changes to architecture, data, or training dynamics can open new scaling axes.
Selected papers
ZAYA1-8B Technical Report (2026)
Robert Washbourne, Rishi Iyer, Tomas Figliolia, Henry Zheng, Ryan Lorig-Roach, Sungyeon Yang, Pritish Yuvraj, Quentin Anthony, Yury Tokpanov, Xiao Yang, Ganesh Nanduru, Stephen Ebert, Praneeth Medepalli, Skyler Szot, Srivatsan Rajagopal, Alex Ong, Bhavana Mehta, Beren Millidge
This technical report presents our state-of-the-art 8B LLM foundation model that we trained in-house end-to-end including pretraining, midtraining, RL. Uses a novel in-house architecture we developed (CCA and Zaya router). ZAYA1-8B outperforms all contemporary models of its size and is competitive with substantially larger models including then-frontier models in certain mathematics and coding tasks.
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design (2025)
Quentin Anthony, Yury Tokpanov, Skyler Szot, Srivatsan Rajagopal, Praneeth Medepalli, Rishi Iyer, Vasu Shyam, Anna Golubeva, Ansh Chaurasia, Xiao Yang, Tomas Figliolia, Robert Washbourne, Drew Thorstensen, Amartey Pearson, Zack Grossbart, Jason van Patten, Emad Barsoum, Zhenyu Gu, Yao Fu, Beren Millidge
A detailed systems paper on our novel full-stack AMD pretraining approach. We are the first to enable large-scale LLM pretraining on an end-to-end AMD stack of MI300x GPUs and AMD Pollara networking.
Scaling Adaptive Depth with Norm-Agnostic Residual Networks (2026)
Tomás Figliolia, Beren Millidge
We introduce norm-agnostic residual streams, a novel method to prevent diminishment of marginal capacity growth with depth which exists in current models.
Can Scale Save Us From Plasticity Loss in Large Language Models? (2026)
J. Fernando Hernandez-Garcia, Tomás Figliolia, Beren Millidge
Here we study whether plasticity loss persists in modern language models and how its onset changes with scale. This thus connects continual learning questions to the current contemporary LLM regime.
The Zamba2 Suite (2024)
Paolo Glorioso, Quentin Anthony, Yury Tokpanov, Anna Golubeva, Vasudev Shyam, James Whittington, Jonathan Pilault, Beren Millidge
Introduces the Zamba2 SSM–Transformer hybrid architecture, combining a Mamba backbone with shared attention to improve efficiency while retaining strong language-model performance. We trained then-SOTA LLMs in the 7B, 3B, and 1B size bracket.
Zyda-2: a 5 Trillion Token High-Quality Dataset (2024)
Yury Tokpanov, Paolo Glorioso, Quentin Anthony, Beren Millidge
An example of the data side of the foundation-model program. We constructed and open-sourced a trillion-token-scale pretraining dataset which outperformed comparable pretraining sets of the time, as well as released the full dataset processing, filtering, and deduplication infrastructure.
Predictive coding and local learning
Much of my PhD and postdoctoral work asked whether powerful learning algorithms such as backpropagation can emerge from local distributed dynamics, and whether predictive coding provides a useful general framework for inference and learning in biological and artificial networks.
Selected papers
Predictive Coding Approximates Backprop along Arbitrary Computation Graphs (2020; later published in Neural Computation)
Beren Millidge, Alexander Tschantz, Christopher L. Buckley
We were the first to demonstrate that local learning algorithms such as predictive coding can approximate backpropagation on arbitrary computation graphs, demonstrating a potential route for backprop-like algorithms to be implemented in neural circuitry.
Inferring neural activity before plasticity as a foundation for learning beyond backpropagation (2024 Nature Neuroscience)
Yuhang Song, Beren Millidge, Tommaso Salvatori, Thomas Lukasiewicz, Zhenghua Xu & Rafal Bogacz
We developed prospective configuration a novel learning algorithm building upon predictive coding and show that it outperforms backpropagation on online and continual learning tasks.
A Theoretical Framework for Inference and Learning in Predictive Coding Networks (2022; ICLR 2023)
Beren Millidge, Yuhang Song, Tommaso Salvatori, Thomas Lukasiewicz, Rafal Bogacz
We developed a general framework for understanding how predictive coding networks differ from backpropagation trained networks, and how predictive coding relates to Gauss-Newton, Target-Propagation and other learning algorithms.
Backpropagation at the Infinitesimal Inference Limit of Energy-Based Models (2022; ICLR 2023)
Beren Millidge, Yuhang Song, Tommaso Salvatori, Thomas Lukasiewicz, Rafal Bogacz
We developed a mathematical framework through which we can understand essentially the entire literature of biological learning algorithms approximating backprop through a unifying abstraction of the infinitesimal inference limit.
Hybrid Predictive Coding: Inferring, Fast and Slow (PNAS 2023)
Alexander Tschantz, Beren Millidge, Anil Seth, Christopher Buckley
We combined iterative predictive-coding inference with amortized inference, thus linking biologically motivated local computation with learned feedforward inference.
Control, active inference, and value learning
An earlier strand of my work studied control and reinforcement learning through the lens of probabilistic inference. I was particularly interested in where exploration terms come from, the relationship between iterative planning and amortized policies, and how agents can flexibly represent and revalue multiple rewards.
Selected papers
Whence the Expected Free Energy? (2020; Neural Computation, 2021)
Beren Millidge, Alexander Tschantz, Christopher Buckley
We analyzed the mathematical origin of expected free energy and the relationship between active-inference objectives and information-seeking exploration.
Reward Bases: A Simple Mechanism for Adaptive Acquisition of Multiple Reward Types (PLOS Computational Biology 2024)
Beren Millidge, Yuhang Song, Armin Lak, Mark E. Walton, Rafal Bogacz
Here, we developed a mechanism for rapidly recombining learned reward components as motivational state changes, enabling ‘zero-shot’ transfer of value functions and behaviour to novel reward functions depending on physiological state.
Deep Active Inference as Variational Policy Gradients (2020; Journal of Mathematical Psychology) Beren Millidge
Here, I was the first to ‘scale up’ active inference to contemporary RL environments and agent scales, demonstrating that active-inference-inspired agents outperformed standard policy gradient and Q-learning approaches in deep RL tasks.
Understanding the Origins of Information-Seeking Exploration in Probabilistic Objectives for Control (2021)
Beren Millidge, Alexander Tschantz, Anil Seth, Christopher Buckley
Here we derive the origin of information-seeking objectives in reinforcement learning as deriving from divergence minimizing rather than reward maximizing functionals.
Associative memory and representations
A smaller but recurring line of my research studies associative memory as a general computational primitive and its connections to representation learning, attention, and graph-based retrieval.
Selected papers
Universal Hopfield Networks: A General Framework for Single-Shot Associative Memory Models (2022; ICML 2022)
Beren Millidge, Tommaso Salvatori, Yuhang Song, Thomas Lukasiewicz, Rafal Bogacz
Here, we unified a large literature of existing disparate associative memory models and placed them all into a common framework by decomposing retrieval into similarity, separation, and projection operations. We demonstrated that this unification allows the immediate implementation of novel similarity and separation functions that outperformed existing associative memory methods.
Associative Memories in the Feature Space (2023; ECAI 2023)
Tommaso Salvatori, Beren Millidge, Yuhang Song, Rafal Bogacz
We demonstrate that associative memory models can be made substantially more efficient and performant if the associative operation is performed upon a learnt latent feature space rather than in raw input/output space. We demonstrate that such latent associative memories outperform then-current methods operating on the output space.
Hybrid Associative Memories (2026)
Leon Lufkin, Tomas Figliolia, Beren Millidge, Kamesh Krishnamurthy
We developed a novel hybrid memory which combined SSMs and full attention in a novel way by using the SSM to process the majority of the sequence while only passing to attention the tokens which are surprising to the SSM. We demonstrated that this outperformed existing SSM hybrid methods.
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG (2024)
Nicholas Alonso, Beren Millidge
We developed a novel page-rank inspired RAG mechanism which allowed perfect and SOTA performance on challenging retrieval benchmarks on context lengths of up to a billion tokens while running in real-time entirely on the CPU
All publications
Below is a complete publication record presented in reverse chronological order. I will try to keep this list up to date, however an always up-to-date list can be found at my Google Scholar.
2026
PUFFER: Incremental Fuzzy Deduplication for Continuously Evolving Corpora (2026)
Xiao Yang, Erik Edward Aldape, Beren Millidge
paper
ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution (2026)
Christopher Warner, Jonas Mago, JR Huml, Beren Millidge
paper
ZONOS2 Technical Report (2026)
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close, Beren Millidge
paper
Can Scale Save Us From Plasticity Loss in Large Language Models? (2026)
J. Fernando Hernandez-Garcia, Tomás Figliolia, Beren Millidge
paper
Scaling Adaptive Depth with Norm-Agnostic Residual Networks (2026)
Tomás Figliolia, Beren Millidge
paper
Zamba2-VL Technical Report (2026)
Hassan Shapourian, Kasra Hejazi, Olabode M. Sule, Beren Millidge
paper
ZAYA1-VL-8B Technical Report (2026)
Hassan Shapourian, Kasra Hejazi, Olabode M. Sule, Beren Millidge
paper
ZAYA1-8B Technical Report (2026)
Robert Washbourne, Rishi Iyer, Tomas Figliolia, Henry Zheng, Ryan Lorig-Roach, Sungyeon Yang, Pritish Yuvraj, Quentin Anthony, Yury Tokpanov, Xiao Yang, Ganesh Nanduru, Stephen Ebert, Praneeth Medepalli, Skyler Szot, Srivatsan Rajagopal, Alex Ong, Bhavana Mehta, Beren Millidge
paper
Hybrid Associative Memories (2026)
Leon Lufkin, Tomas Figliolia, Beren Millidge, Kamesh Krishnamurthy
paper
ZUNA: Flexible EEG Superresolution with Position-Aware Diffusion Autoencoders (2026)
Christopher Warner, Jonas Mago, JR Huml, Mohamed Osman, Beren Millidge
paper | code
Online Vector Quantized Attention (2026)
Nick Alonso, Tomas Figliolia, Beren Millidge
paper
2025
Equivalence of Personalized PageRank and Successor Representations (2025)
Beren Millidge
paper
Generalizing E-prop to Deep Networks (2025)
Beren Millidge
paper
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design (2025)
Quentin Anthony, Yury Tokpanov, Skyler Szot, Srivatsan Rajagopal, Praneeth Medepalli, Anna Golubeva, Vasu Shyam, Robert Washbourne, Rishi Iyer, Ansh Chaurasia, Tomas Figliolia, Xiao Yang, Drew Thorstensen, Amartey Pearson, Zack Grossbart, Jason van Patten, Emad Barsoum, Zhenyu Gu, Yao Fu, Beren Millidge
paper
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space (2025)
Tomas Figliolia, Nicholas Alonso, Rishi Iyer, Quentin Anthony, Beren Millidge
paper
2024
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG (2024)
Nicholas Alonso, Beren Millidge
paper
The Zamba2 Suite: Technical Report (2024)
Paolo Glorioso, Quentin Anthony, Yury Tokpanov, Anna Golubeva, Vasudev Shyam, James Whittington, Jonathan Pilault, Beren Millidge
paper | code
Reward Bases: A simple mechanism for adaptive acquisition of multiple reward types (2024)
Beren Millidge, Yuhang Song, Armin Lak, Mark E Walton, Rafal Bogacz
paper | code
Zyda-2: a 5 Trillion Token High-Quality Dataset (2024)
Yury Tokpanov, Paolo Glorioso, Quentin Anthony, Beren Millidge
paper
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters (2024)
Vasudev Shyam, Jonathan Pilault, Emily Shepperd, Quentin Anthony, Beren Millidge
paper | code
Zyda: A 1.3 T Dataset for Open Language Modeling (2024)
Yury Tokpanov*, Beren Millidge*, Paolo Glorioso, Jonathan Pilault, Adam Ibrahim, James Whittington, Quentin Anthony
paper | code
Toward Conversational Agents with Context and Time Sensitive Long-term Memory (2024)
Nicholas Alonso, Tomas Figliolia, Anthony Ndirango, Beren Millidge
paper
Zamba: A Compact SSM Hybrid Model (2024)
Paolo Glorioso*, Quentin Anthony*, Yury Tokpanov*, James Whittington, Jonathan Pilault, Adam Ibrahim, Beren Millidge*
paper | code
A Review of Neuroscience-Inspired Machine Learning (2024)
Alexander Ororbia, Ankur Mali, Adam Kohan, Beren Millidge, Tommaso Salvatori
paper
Natural Induction: Spontaneous adaptive organisation without natural selection (2024)
Christopher L Buckley, Tim Lewens, Michael Levin, Beren Millidge, Alec Tschantz, Richard A Watson
paper
BlackMamba: Mixture of Experts for State Space Models (2024)
Quentin Anthony*, Yury Tokpanov*, Paolo Glorioso*, Beren Millidge*
paper | code
2023
Collective Behaviour from Surprise Minimization (2023)
Conor Heins, Beren Millidge, Lancelot Da Costa, Richard Mann, Karl Friston, Iain Couzin
paper
Predictive Coding Networks for Temporal Prediction (2023)
Beren Millidge, Mufeng Tang, Mahyar Osanlouy, Rafal Bogacz
paper
Exploring Action-Centric Representations through the Lens of Rate-Distortion Theory (2023)
Miguel De Llanza Varona, Christopher Buckley, Beren Millidge
paper
Causal Inference via Predictive Coding (2023)
Tommaso Salvatori, Luca Pinchetti, Amine M’Charrak, Beren Millidge, Thomas Lukasiewicz
paper
Associative Memories in the Feature Space (2023)
Tommaso Salvatori, Beren Millidge, Yuhang Song, Rafal Bogacz
paper
From the free energy principle to a confederation of Bayesian mechanics. Reply to comments on” How particular is the physics of the free energy principle?” (2023)
Miguel Aguilera, Beren Millidge, Alexander Tschantz, Christopher Buckley
paper
2022
Interpreting Neural Networks through the Polytope Lens (2022)
Sid Black, Lee Sharkey, Leo Grinsztajn, Eric Winsor, Dan Braun, Jacob Merizian, Kip Parker, Carlos Ramón Guevara, Beren Millidge, Gabriel Alfour, Connor Leahy
paper
Generalized Predictive Coding: Bayesian Inference in Static and Dynamic models (2022)
Andre Ofner, Beren Millidge, Sebastian Stober
paper
Recurrent predictive coding models for associative memory employing covariance learning (2022)
Mufeng Tang, Tommaso Salvatori, Beren Millidge, Yuhang Song, Thomas Lukasiewicz, Rafal Bogacz
paper
Incremental Predictive Coding: A Parallel and Fully Automatic Learning Algorithm (2022)
Tommaso Salvatori, Yuhang Song, Beren Millidge, Zhenghua Xu, Lei Sha, Cornelius Emde, Rafal Bogacz, Thomas Lukasiewicz
paper
Predictive Coding Beyond Gaussian Distributions (2022)
Luca Pinchetti, Tommaso Salvatori, Yordan Yordanov, Beren Millidge, Yuhang Song, Thomas Lukasiewicz
paper
Designing ecosystems of intelligence from first principles (2022)
Karl J Friston, Maxwell JD Ramstead, Alex B Kiefer, Alexander Tschantz, Christopher L Buckley, Mahault Albarracin, Riddhi J Pitliya, Conor Heins, Brennan Klein, Beren Millidge, Dalton AR Sakthivadivel, Toby St Clere Smithe, Magnus Koudahl, Safae Essafi Tremblay, Capm Petersen, Kaiser Fung, Jason G Fox, Steven Swanson, Dan Mapes, and Gabriel René
paper
Capsule Networks as Generative Models (2022)
Alex B Kiefer*, Beren Millidge*, Alexander Tschantz*, Christopher Buckley
paper | Alex’s code, my code
Preventing Deterioration of Classification Accuracy in Predictive Coding Networks (2022)
Paul F Kinghorn, Beren Millidge, Christopher L Buckley
paper
A Theoretical Framework for Inference and Learning in Predictive Coding Networks (2022)
Beren Millidge, Yuhang Song, Tommaso Salvatori, Thomas Lukasiewicz, Rafal Bogacz
paper | code
Successor Representation Active Inference (2022)
Beren Millidge, Christopher L Buckley
paper | code
A Theoretical Framework for Inference Learning (2022)
Nick Alonso, Beren Millidge, Jeff Krichmar, Emre Neftci
paper | code
Backpropagation at the Infinitesimal Inference Limit of Energy-Based Models: Unifying Predictive Coding, Equilibrium Propagation, and Contrastive Hebbian Learning (2022)
Beren Millidge, Yuhang Song, Tommaso Salvatori, Thomas Lukasiewicz, Rafal Bogacz
paper | code
On Bayesian Mechanics: A physics of and by beliefs (2022)
Maxwell JD Ramstead, Dalton AR Sakthivadivel, Conor Heins, Magnus Koudahl, Beren Millidge, Lancelot Da Costa, Brennan Klein, Karl J Friston
paper
Inferring Neural Activity Before Plasticity: A Foundation for Learning Beyond Backpropagation (2022)
Yuhang Song, Beren Millidge, Tommaso Salvatori, Thomas Lukasiewicz, Zhenghua Xu, Rafal Bogacz
paper | code
Reward Bases: Instantaneous Reward Revaluation with Temporal Difference Learning (2022)
Beren Millidge, Mark Walton, Rafal Bogacz
paper | code
Hybrid Predictive Coding: Inferring, Fast and Slow (2022)
Alexander Tschantz*, Beren Millidge*, Anil Seth, Christopher Buckley.
paper | code
Predictive Coding: Towards a Future of Deep Learning beyond Backpropagation? (2022)
Beren Millidge*, Tommaso Salvatori*, Yuhang Song, Rafal Bogacz, Thomas Lukasiewicz
paper
Universal Hopfield Networks: A General Framework for Single-Shot Associative Memory Models (2022)
Beren Millidge, Tommaso Salvatori, Yuhang Song, Thomas Lukasiewicz, Rafal Bogacz
paper | code
Learning on Arbitrary Graph Topologies via Predictive Coding (2022)
Tommaso Salvatori, Luca Pinchetti, Beren Millidge, Yuhang Song, Rafal Bogacz, Thomas Lukasiewicz
paper
pymdp: A Python library for active inference in discrete state spaces (2022)
Conor Heins, Beren Millidge, Daphne Demekas, Brennan Klein, Karl Friston, Iain Couzin, Alexander Tschantz
paper | code
2021
Active Inference in Robotics and Artificial Agents: Survey and Challenges (2021)
Pablo Lanillos, Cristian Meo, Corrado Pezzato, Ajith Anil Meera, Mohamed Baioumy, Wataru Ohata, Alexander Tschantz, Beren Millidge, Martijn Wisse, Christopher L. Buckley, Jun Tani
paper
Habitual and Reflective Control in Hierarchical Predictive Coding (2021)
Paul F Kinghorn, Beren Millidge, Christopher L Buckley
paper
A Mathematical Walkthrough and Discussion of the Free Energy Principle (2021)
Beren Millidge, Anil Seth, Christopher Buckley
paper
Predictive Coding: A Theoretical and Experimental Review (2021)
Beren Millidge, Anil Seth, Christopher Buckley
paper
Applications of the Free Energy Principle to Machine Learning and Neuroscience (2021)
Beren Millidge
paper | code
Online Reinforcement Learning with Sparse Rewards through an Active Inference Capsule (2021)
Alejandro Daniel Noel, Charel van Hoof, Beren Millidge
paper | code
Towards a Mathematical Theory of Abstraction (2021)
Beren Millidge
paper
How Particular is the Physics of the Free Energy Principle (2021)
Miguel Aguilera, Beren Millidge, Alexander Tschantz, Christopher Buckley
paper
Understanding the Origins of Information-Seeking Exploration in Probabilistic Objectives for Control (2021)
Beren Millidge, Alexander Tschantz, Anil Seth, Christopher Buckley
paper | code
Neural Kalman Filtering (2021)
Beren Millidge, Alexander Tschantz, Anil Seth, Christopher Buckley
paper | code
2020
Sophisticated Active Inference: Simulating Anticipatory Affective Dynamics of Imagining Future Events (2020)
Casper Hesp, Alexander Tschantz, Beren Millidge, Maxwell Ramstead, Karl Friston, Ryan Smith
paper
Published in IWAI IEEE workshop on Active Inference
Investigating the Scalability and Biological-Plausibility of the Activation Relaxation Algorithm (2020)
Beren Millidge, Alexander Tschantz, Anil Seth, Christopher L Buckley
paper | code
Published in NeurIPS 2020 workshop on Backpropagation in the Brain
Relaxing the Constraints on Predictive Coding Models (2020)
Beren Millidge, Alexander Tschantz, Anil Seth, Christopher L Buckley
paper | code
Activation Relaxation: A Local Dynamical Approximation to Backprop in the Brain (2020)
Beren Millidge, Alexander Tschantz, Anil Seth, Christopher L Buckley
paper | code
The Acquisition of Culturally Patterned Attention Styles under Active Inference (2020)
Axel Constant, Alexander Tschantz, Beren Millidge, Felipe Criado-Boado, Luis M Martinez, Johannes Müller, Andy Clark.
paper | code
Control as Hybrid Inference (2020)
Alexander Tschantz, Beren Millidge, Anil Seth, Christopher Buckley
paper
Published in ICML 2020 workshop.
On the Relationship between Control as Inference and Active Inference (2020)
Beren Millidge, Alexander Tschantz, Anil Seth, Christopher Buckley
paper
Published in IWAI IEEE workshop on Active Inference
Reinforcement Learning as Iterative and Amortised Inference (2020)
Beren Millidge*, Alexander Tschantz*, Christopher Buckley
paper
Predictive Coding Approximates Backprop Along Arbitrary Computation Graphs (2020)
Beren Millidge, Alexander Tschantz, Christopher Buckley
paper | code
Curious Inferences: Reply to Sun & Firestone on the Dark Room Problem (2020)
Anil Seth, Beren Millidge, Christopher Buckley, Alexander Tschantz
paper
Published in Trends in Cognitive Science
Whence the Expected Free Energy (2020)
Beren Millidge, Alexander Tschantz, Christopher Buckley
paper
Published in Neural Computation
Reinforcement Learning Through Active Inference (2020)
Alexander Tschantz*, Beren Millidge*, Anil Seth, Christopher Buckley
paper | code
Published in Bridging AI and Cognitive Science (ICLR 2020) workshop
2019
Deep Active Inference as Variational Policy Gradients (2019)
Beren Millidge
Published in Journal of Mathematical Psychology
paper | code
Combining Active Inference and Hierarchical Predictive Coding: A Tutorial Introduction and Case-Study (2019)
Beren Millidge
paper | code
Implementing Predictive Processing and Active Inference: Preliminary Steps and Results (2019)
Beren Millidge
paper | code
Vocal imitation can create acoustic attractors to guide mothers to pups in a crowded colony of Mexican free-tailed bats: A case-study of computational modelling in behavioural biology (2019)
Richard Shillcock, Beren Millidge, Andrea Ravignani
paper | code
Fixational Eye Movements: Data Augmentation for the Brain? (2019)
Beren Millidge
paper | code
Exploring infant vocal imitation in Tadarida brasiliensis mexicana (2019)
Richard Shillcock, Beren Millidge, Andrea Ravignani (2019)
Published in Neurobiology of Speech and Language
paper
2018
A Predictive Processing Account of Bottom-Up Visual Saliency Using Cross-Predicting Autoencoders (2018)
Beren Millidge, Richard Shillcock
paper | code