Token-wise residual latent adapters: Steering seq2Seq models for protein fitness extrapolation
2026
Protein design requires extrapolating beyond training data to achieve higher fitness. State-of-the-art methods typically fine-tune billion-parameter language models end-to-end, often combined with external scorers, data distillation, and multiple rounds of iterative refinement. We introduce a residual latent adapter, a 5M parameter MLP inserted between the encoder and decoder of a frozen ProtT5-3B model, which learns a token-wise residual transformation on encoder embeddings via a simple MSE objective. In a single forward pass with no external scorer, RLA achieves strong extrapolation rates and top-100 candidate fitness relative to methods requiring 600 more trainable parameters and multi-stage pipelines, with the largest gains on the harder extrapolation benchmarks. Our results demonstrate that a compact residual transformation in latent space provides a simple, data-efficient, and compute-efficient approach to in-silico protein fitness extrapolation.
Research areas