Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Sciences

1ETH AI Center, 2IBM Research Europe, 3 SAM ETH Zurich, 4SDSC
NeurIPS 2026

Disentangled Latent Control: Phaedra factorizes fields into independent $z_\mu$ (Morphology) and $z_\alpha$ (Amplitude) tokens. By combining, for example, the local vortices from Sample B ($z_{\mu,B}$) with the global magnitude profile of Sample A ($z_{\alpha,A}$), the reconstruction maintains the fine-scale turbulent structure of one while strictly obeying the physical dynamic range of the other.

Abstract

Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned. As existing tokenizers are designed for realistic visual perception, we investigate whether these are optimal for scientific images, which exhibit a large dynamic range and require token embeddings to retain physical properties. We propose Phaedra, inspired by classical shape-gain quantization and proper orthogonal decomposition. We demonstrate that Phaedra consistently improves reconstruction across a range of PDE datasets and shows strong out-of-distribution generalization to unknown PDEs and real-world Earth observation data.

Phaedra Comparisons

Phaedra consistently improves reconstruction across PDE datasets, capturing fine details and precise magnitudes critical for scientific simulation. This is achieved by splitting the embeddings into shape and gain components, allowing for high-fidelity tokenization that retains physical properties, even in out-of-distribution scenarios.

Data distributions

Data distributions after normalization. Natural images have a fixed range and uniform distribution, even after normalization. Physical datasets, however, have a much larger range of values with outliers far outside the nominal range.

Methodology

Phaedra leverages a disentangled hierarchical representation inspired by shape-gain quantization. The tokenizer factorizes latent embeddings into two distinct components: morphology ($z_\mu$), which captures spatial structures and topological features via Finite Scalar Quantization (FSQ), and amplitude ($z_\alpha$), which preserves the dynamic range and physical magnitudes. By incorporating an approximately continuous channel for amplitude, Phaedra avoids the precision loss common in standard discrete codebooks. This architecture effectively mirrors Proper Orthogonal Decomposition (POD) by separating spatial modes from their scalar coefficients, ensuring that physical properties such as conservation laws and sharp gradients are maintained even at high compression ratios.

Phaedra Pipeline

The Phaedra pipeline uses a single encoder to compute two sets of embeddings; the shape embeddings are quantized using sparse, high-dimensional FSQ to capture spatial structures, while the gain embeddings are passed through a dense, 1-dimensional quantizer which approximates a continuous channel, preserving physical magnitudes. Following quantization, a learned recombination operator combines these two embeddings and passes them to the decoder for the final reconstruction.

Interactive Latent Space Exploration

Use the sliders below to independently interpolate the morphology ($z_\mu$) and amplitude ($z_\alpha$) tokens between two timesteps of the Kelvin-Helmholtz (KH) instability dataset. This example illustrates the disentanglement of features via the amplitude-morphology latent representations. Interpolating between the amplitude tokens shows larger changes within the overall structure of the flow, while changing morphology tokens results in finer changes to the turbulent structures, while maintaining the same overall dynamic range. By combining the morphology of one sample with the amplitude of another, we can generate reconstructions that maintain the fine-scale turbulent structure of one sample while strictly obeying the physical dynamic range of the other. Please note, this example is only for illustrative purposes and does not represent a true interpolation in the latent space, as the model was not trained with an explicit disentanglement loss.

KH t=0.95 (A) KH t=1.00 (B)
KH t=0.95 (A) KH t=1.00 (B)
Active Reconstruction

Active Reconstruction

PDE Reconstruction Benchmarks

All models are trained on the same set of data, comprising Compressible Euler time-series defined by 4 different classes of initial conditions and Incompressible Navier-Stokes defined by 2 different classes of initial conditions. While we present some competitive tokenizers below, a comprehensive comparison against additional baselines is detailed in the main paper. We evaluate Phaedra across three levels of difficulty.

  • ID (In-Distribution): Test samples from the same PDE family and parameters as training.
  • OD1 (Out-of-Distribution): Shifts in the PDE coefficients and initial conditions.
  • OD2 (Out-of-Distribution): Shifts in the PDEs that define the problem.
ID: CEU Curved Riemann
OD1: Airfoil
OD2: Acoustic Wave
ID Input
Input
OD1 Input
Input
OD2 Input
Input
ID Recon
Phaedra Reconstruction
OD1 Recon
Phaedra Reconstruction
OD2 Recon
Phaedra Reconstruction
Model Dataset nMAE ↓ nRMSE ↓ Δσ²loc ↓ γmin ↑
VQ-VAE-2ID3.0245.06915.0279.1%
OD12.1135.85421.5587.5%
OD24.4495.83317.0668.9%
FSQID2.6034.29211.2985.3%
OD11.8764.31420.5593.8%
OD23.8314.99711.6568.0%
Phaedra (Ours) ID1.5222.4895.9693.6%
OD11.2172.4426.4798.0%
OD22.5003.3635.8279.9%

Sentinel-2 L1C Earth Observation Data

Zero-Shot Evaluation: All models were evaluated without fine-tuning on Earth Observation data, relying on representations learned from synthetic physical simulations.
Group Model rL₁ ↓ rL₂ ↓ Δσ²loc ↓ γmin ↑
Baseline Continuous 7.426 8.100 62.61 90.02%
Compression: 4$\times$4 Phaedra₄ (Ours) 8.895 9.749 128.5 93.17%
FSQ₄ 11.053 12.405 215.8 77.62%
Compression: 8$\times$8 Phaedra₈ (Ours) 9.900 11.475 163.9 62.72%
Cosmos₈ 16.717 19.245 19,566 79.70%

Metrics calculated on native 13-band resolutions (10m, 20m, 60m).

Sentinel-2 Reconstruction results
Figure 1: Comparison of Ground Truth vs. Phaedra Reconstruction for Sentinel-2 L1C imagery.

Neural Operators on Phaedra Tokens

Tokens are only useful if models can learn on them. We train a 38M-parameter encoder–decoder transformer that maps the tokens of an initial condition, together with a lead time $t$, to the tokens at time $t$; the frozen Phaedra decoder turns the predicted tokens back into density, velocity and pressure. The same transformer is trained on FSQ and VQ-VAE-2 tokens and on the latents of a continuous autoencoder, and we compare with FNO, CNO and a ViT that work directly on the fields — every model has about 38M parameters (44M for the larger VQ-VAE-2 vocabularies) and is trained separately on three compressible Euler datasets from Poseidon: Kelvin–Helmholtz (KH), curved Riemann (RC) and Riemann–Kelvin–Helmholtz (RKH), which the tokenizer never saw during its training.

Predicted tokens and the physics decoded from them

Variable:

The first test trajectory of each dataset, from $t = 0$ to the final time $t = 0.7$. From left to right: the ground truth; the amplitude and morphology tokens predicted by the transformer on the 32×32 token grid (morphology tokens are coloured by their FSQ code); the field decoded from these tokens; and the absolute error, labelled with the relative $L_1$ error of the frame. Every frame is predicted in a single step from the tokens of the initial condition; the first frame shows the tokenized initial condition and its reconstruction. The video is rendered from the released weights (scripts/website/render_token_video.py), and its errors are identical to those in the evaluation below.

Accuracy at the final time

Relative $L_1$ error (%) at $t = 0.7$, averaged over $\rho, u, v, p$ and the 240 test trajectories of each dataset, ± the 95% bootstrap confidence interval. Every model is shown with its best prediction strategy (tag: d direct $0 \to t$, 2s autoregressive 2-step, 662 rollout $0 \to 0.3 \to 0.6 \to 0.7$); bold marks the best model per dataset. Phaedra tokens are the best of the discrete representations on all three datasets, with 16%, 28% and 43% lower error than FSQ tokens on KH, RC and RKH. They give the lowest error of all models on RC and RKH; on KH the ViT (9.18%) and the continuous-latent transformer (9.28%) are slightly more accurate than Phaedra (9.50%).

ModelInputKHRCRKH
Phaedra transformerPhaedra tokens9.50 ± 0.34 d27.21 ± 0.71 66210.23 ± 0.80 662
FSQ transformerFSQ tokens11.28 ± 0.34 66238.05 ± 1.05 66217.99 ± 1.37 662
VQ-VAE-2 transformerVQ-VAE-2 tokens15.14 ± 0.48 66242.40 ± 0.63 d24.52 ± 1.15 662
Continuous transformercontinuous AE latents9.28 ± 0.30 662119.05 ± 2.96 d11.56 ± 0.70 662
FNOphysical fields12.55 ± 0.36 66235.63 ± 0.69 66216.66 ± 0.91 662
CNOphysical fields10.15 ± 0.32 2s29.92 ± 0.74 2s11.82 ± 0.82 662
ViTphysical fields9.18 ± 0.27 66239.68 ± 0.58 d13.20 ± 0.79 662
Per-variable errors ($\rho, u, v, p$) of all 21 models
ModelDataρuvpAverageStrategy
PhaedraKH4.788.4324.290.519.50 ± 0.34direct
PhaedraRC20.6640.7340.427.0227.21 ± 0.716-6-2
PhaedraRKH6.5216.0115.642.7410.23 ± 0.806-6-2
FSQKH5.3010.9128.270.6211.28 ± 0.346-6-2
FSQRC28.1257.0856.7710.2138.05 ± 1.056-6-2
FSQRKH10.3727.7328.785.0817.99 ± 1.376-6-2
VQ-VAE-2KH7.9513.9038.040.6815.14 ± 0.486-6-2
VQ-VAE-2RC29.5463.1764.3912.4842.40 ± 0.63direct
VQ-VAE-2RKH14.7338.0937.297.9724.52 ± 1.156-6-2
ContinuousKH4.648.3423.530.619.28 ± 0.306-6-2
ContinuousRC77.95197.31156.9344.00119.05 ± 2.96direct
ContinuousRKH7.0918.0917.873.1811.56 ± 0.706-6-2
FNOKH6.2411.1832.230.5412.55 ± 0.366-6-2
FNORC26.3153.6253.429.1635.63 ± 0.696-6-2
FNORKH10.0826.3625.674.5416.66 ± 0.916-6-2
CNOKH5.069.1025.930.5110.15 ± 0.322-step
CNORC22.6544.4444.667.9429.92 ± 0.742-step
CNORKH6.7918.8318.503.1711.82 ± 0.826-6-2
ViTKH4.648.2523.330.519.18 ± 0.276-6-2
ViTRC28.7660.1759.5110.2839.68 ± 0.58direct
ViTRKH9.0920.8619.663.1813.20 ± 0.796-6-2

All tables as CSV, Markdown and LaTeX: docs/results. The released RKH operators are the checkpoints evaluated in the paper: step 56k of the Phaedra transformer and epoch 59 of the continuous-latent transformer, whose training diverged afterwards.

Error over time

Relative $L_1$ error over the prediction horizon (mean over $\rho, u, v, p$ and the 240 test trajectories). The snapshots are $\Delta t = 0.05$ apart. Direct predicts every time in a single step from $t = 0$, 2-step chains predictions two snapshots ahead ($2\Delta t = 0.1$ per step, $0 \to 0.1 \to \dots \to 0.7$), and 6-6-2 rolls out $0 \to 0.3 \to 0.6 \to 0.7$. Click a model in the legend to hide it.

Dataset:
Strategy:

All models at $t = 0.7$

Ground truth and every operator at t = 0.7 (KH, first test trajectory)
The first test trajectory at the final time, all four variables. Each model is shown with the strategy of the table above; the label gives its average relative $L_1$ error on this trajectory. The colour scales follow the ground truth.

The tokenizer sets the floor

Decoding the ground-truth tokens of the test data gives the lowest error that any operator on a representation can reach. Phaedra's floor is 1.6–1.9× lower than that of FSQ and 2.3–2.5× lower than that of VQ-VAE-2 on the same 32×32 token grid; only the continuous autoencoder, which is not discrete, is lower. Relative $L_1$ error (%) at $t = 0.7$, mean over $\rho, u, v, p$.

TokenizerKHRCRKH
Phaedra0.543.301.49
FSQ1.015.162.36
VQ-VAE-21.378.393.45
Continuous AE (not discrete)0.271.480.65

Masked Autoencoding on Token Grids

Phaedra tokens also support self-supervised learning. A masked autoencoder (a transformer with about 48M parameters) sees a token grid with 75% of its tokens hidden and reconstructs them. It is pre-trained on KH, RC and RKH together and then fine-tuned on each dataset. The table gives the relative $L_1$ error (%) of the fields decoded from the reconstructed token grids, mean over $\rho, u, v, p$ (number of test trajectories in parentheses).

Tokenspre-trained on 3 PDEs
evaluated on RKH
fine-tuned
evaluated on KH
fine-tuned
evaluated on RC
fine-tuned
evaluated on RKH
Phaedra16.43 (10 traj.)4.76 (10 traj.)22.17 (10 traj.)7.27 (10 traj.)
FSQ16.91 (10 traj.)6.29 (1 traj.)28.20 (1 traj.)7.07 (1 traj.)

The masks are random, so each number comes from one evaluation run. The fine-tuned FSQ models were evaluated on a single test trajectory.

Code, Models and Data

Code

The tokenizer (also as a small pip-installable package), token generation, the 21 neural operators, masked autoencoding, and the evaluation that reproduces every number on this page. MIT license.

🤗 Models

Four tokenizers (Phaedra, FSQ, VQ-VAE-2, continuous AE), the 21 neural operators and eight masked autoencoders as safetensors, each with its training configuration. CC BY-NC 4.0.

🤗 Tokenized data

Phaedra, FSQ and VQ-VAE-2 tokens and continuous latents of KH, RC and RKH: 10,000 trajectories with 21 time steps each (44 GB). CC BY-NC 4.0, following the Poseidon source data.

Tokenize a field

git clone https://github.com/camlab-ethz/Phaedra.git && cd Phaedra && pip install -e .

import torch
from phaedra import load_pretrained

model = load_pretrained()                        # the paper's tokenizer, downloaded from Hugging Face
x = torch.randn(1, 1, 128, 128)                  # a field normalized per variable: (x - mean) / std
quant, _, (morph, amp), _ = model.encode(x)      # morphology and amplitude tokens on the 32x32 grid
x_rec = model.decode(quant)

Reproducing the operator and MAE results, training from scratch and preparing the Poseidon data are described in the README.

Conclusion & Discussion

Our findings demonstrate that Phaedra successfully closes the "fidelity gap" that typically limits discrete tokenization in scientific applications. By balancing structural discretization with magnitude preservation, the model achieves superior reconstruction across diverse PDE families and exhibits remarkable zero-shot transfer to complex Earth observation data. The higher fidelity carries over to learning on the tokens: neural operators trained on Phaedra tokens are more accurate than those trained on the other discrete tokenizers. This suggests that the shape-gain prior serves as a fundamental inductive bias for physical systems, enabling models to generalize across scales and disciplines. As we move toward larger scientific foundation models, Phaedra provides a robust framework for transforming continuous physical fields into discrete sequences without sacrificing the precision required for rigorous scientific analysis.

BibTeX

@inproceedings{lingsch2026phaedra,
  title         = {Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Sciences},
  author        = {Lingsch, Levi and Kissas, Georgios and Jakubik, Johannes and Mishra, Siddhartha},
  booktitle     = {Advances in Neural Information Processing Systems},
  year          = {2026},
  eprint        = {2602.03915},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2602.03915}
}