NIVA Brings A Multimodal Foundation Model for into Focus

This paper introduces NIVA, a new AI model that helps improve weather and climate predictions by understanding how different parts of the Earth system interact. It focuses on the ocean and atmosphere to make better forecasts beyond two weeks.

Analyze with PDFdigest

Content & Liability Disclaimer

This article and its accompanying video are automated summaries derived from NIVA: A Multimodal Foundation Model for Actionable Earth System Intelligence by Anisha Pal, Aodhan Sweeney, Kyle Heyblom, Kalai Ramea. The original research was conducted solely by the paper's authors; PDFdigest did not conduct any of the research and makes no claims of ownership over the underlying scientific work.

The video narration is generated by artificial intelligence and references the paper's authors for attribution. The video is not narrated by any of the paper's authors. This content may contain inaccuracies, omissions, or misinterpretations of the original research. First-person language (e.g., "we found", "our results") reflects the original authors' voice, not PDFdigest's. Always read the original paper for accurate, verified information before making any decisions based on this content.

This content is provided "as is" without any warranties, express or implied. Simulated systems OÜ, its officers, directors, employees, and agents shall not be liable for any direct, indirect, incidental, special, consequential, or punitive damages arising from your use of, reliance on, or access to this content, including but not limited to errors, omissions, or misinterpretations of the original research. This disclaimer applies to the fullest extent permitted by applicable law.

Key Takeaways
  1. 1 To complement the contrastive objective, we incorporate a lightweight auxiliary reconstruction loss on the atmospheric modality, where the input field is reconstructed from its latent representation.
  2. 2 The reconstruction term is weighted by α ∈ [0, 0.1] to provide a stabilizing signal without dominating the primary objective.
  3. 3 The similarity metric (cosine similarity) and corresponding loss formulation for this contrastive objective is detailed in Sec.
  4. 4 Our pretraining framework follows a contrastive learning objective with separate encoders for oceanic and atmospheric states.

Introduction

Foundation models have emerged as a powerful paradigm for learning general-purpose representations from largescale data, enabling flexible adaptation across a wide range of downstream tasks. A growing line of work seeks to develop foundation models in the context of Earth system science, where the high-dimensional, non-linear dynamics of the system, combined with increasing availability of long-term datasets, motivate data-driven approaches to learn these processes directly from observations.

However, despite these advances, existing models remain limited in scope, as they primarily focus on atmospheric variables and do not fully capture the coupled dynamics of the broader Earth system, such as the ocean, sea ice, and land surface.

Furthermore, existing models are largely trained on observational reanalysis datasets, which are dominated by atmospheric observations and data assimilation systems.

Important Note

This limitation is critical because predictability at subseasonal-to-seasonal (S2S) timescales depends on interactions between the atmosphere and slower components such as the ocean, sea ice, and land surface .

Important Note

1, while foundation models for weather & climate learn generalized representations supporting a range of downstream tasks, they remain limited to shorter timescales .

Research Question

Our pretraining framework follows a contrastive learning objective with separate encoders for oceanic and atmospheric states. The similarity metric (cosine similarity) and corresponding loss formulation for this contrastive objective is detailed in Sec.

We assume a one-to-one correspondence between modalities within each temporal window and optimize the objective symmetrically in both directions (ocean→atmosphere and atmosphere→ocean).

To complement the contrastive objective, we incorporate a lightweight auxiliary reconstruction loss on the atmospheric modality, where the input field is reconstructed from its latent representation.

Methodology

In domains such as natural language processing and computer vision, these models have shifted the focus from task-specific solutions to unified architectures that capture underlying structure in complex systems . Consequently, while they improve scalability and cross-task generalization, they do not address the core challenge of learning coupled Earth system dynamics.

Study Design

The task is a regression problem that is optimized using Huber loss .

How PDFdigest Helps You Understand Research

Instant Paper Analysis

Get structured summaries and key findings from dense PDFs in seconds.

Visual Explanations

Turn complex methods, figures, and results into clearer visual breakdowns.

AI-Powered Q&A

Ask focused questions and get answers grounded in the paper.

Try PDFdigest Free

Results & Findings

Early Earth system foundation models, including Cli-maX , AtmoRep , Aurora , and Prithvi WxC , demonstrate the promise of learning taskagnostic representations from large-scale atmospheric data. These models support a range of downstream applications, including forecasting, downscaling, and event counterfactuals, often rivaling traditional numerical approaches in shortrange prediction tasks.

  • Early Earth system foundation models, including Cli-maX , AtmoRep , Aurora , and Prithvi WxC , demonstrate the promise of learning taskagnostic representations from large-scale atmospheric.
  • These models support a range of downstream applications, including forecasting, downscaling, and event counterfactuals, often rivaling traditional numerical approaches in shortrange prediction tasks.
  • As a result, restricting their predictive skill to shorter timescales, as atmospheric predictability beyond ≈ 2 weeks depends on memory stored in these slow-evolving components .
  • This leads to an atmosphere-centric view of the Earth system, in which slowly evolving components such as the ocean are under observed and weakly constrained, especially.
  • As a result, by emphasizing high-frequency atmospheric dynamics while underrepresenting lower-frequency variability and coupling, current foundation models struggle to address key challenges such as S2S forecasting.
Important Note

The learned latent space is designed to support transfer to downstream tasks such as climate risk assessment, extreme event prediction, and climate attribution; we leave a systematic evaluation of these capabilities to future work.

Important Note

NIVA addresses this limitation by operating at lower temporal resolutions, enabling the model to better represent slowly evolving subsystems and their interactions.

Practical Applications

This acts as a regularizer, encouraging retention of fine-grained spatial structure that may otherwise be suppressed under purely contrastive training. In addition, structured regions of elevated similarity appear off the diagonal, suggesting that a single ocean state may correspond to multiple dynamically consistent atmospheric realizations.

This metric is well-suited to our setting, as it accounts for the fact that multiple atmospheric states may be physically consistent with a given ocean condition.

Consequently, even when the exact temporal match is not ranked first, closely related atmospheric states may still occupy top positions.

Numerical Earth System Models

Numerical Earth System Models (ESMs) simulate Earth’s climate and weather by integrating physical equations. They are essential for forecasting and climate risk assessment but are computationally expensive and require substantial domain expertise, limiting their accessibility and iterative development.

Multimodal Representation Learning

Multimodal representation learning integrates heterogeneous data sources to derive rich semantic representations. Recent advances in contrastive approaches have shown promise in capturing relationships across modalities, which this work extends to model the coupled dynamics between oceanic and atmospheric states.

Foundation Models for Weather and Climate

While existing foundation models for weather and climate improve scalability, they often treat the Earth system as an atmosphere-only problem. NIVA reformulates this by treating different components as distinct modalities, capturing cross-component dependencies and enabling representations that reflect the coupled nature of the Earth system.

Source Paper Figures and Captions

Source-paper figure
Figure 1. Schematic of the NIVA end-to-end pipeline. The framework uses ESM output to pretrain the foundation model that learns coupled Earth system representations, which can be transferred to a range of downstream tasks.
Figure 1. Schematic of the NIVA end-to-end pipeline. The framework uses ESM output to pretrain the foundation model that learns coupled Earth system representations, which can be transferred to a range of downstream tasks.
Source-paper figure
Figure 2. Pretraining approach. Aggregated oceanic and atmospheric states are processed by separate encoders and jointly optimized with a contrastive objective to learn a shared latent representation of coupled dynamics.
Figure 2. Pretraining approach. Aggregated oceanic and atmospheric states are processed by separate encoders and jointly optimized with a contrastive objective to learn a shared latent representation of coupled dynamics.
Source-paper figure
Figure 3. Pretraining metrics for NIVA on training and validation data (a) Cosine similarity matrix, where blue indicates high similarity and red indicates low similarity. The strong diagonal structure reflects correct alignment between paired ocean and atmospheric states in the latent space. (b) Reciprocal Rank (RR) distribution over 2000 samples. The concentration of samples at RR = 1 indicates that the model reliably retrieves the correct cross-modal pairs.
Figure 3. Pretraining metrics for NIVA on training and validation data (a) Cosine similarity matrix, where blue indicates high similarity and red indicates low similarity. The strong diagonal structure reflects correct alignment between paired ocean and atmospheric states in the latent space. (b) Reciprocal Rank (RR) distribution over 2000 samples. The concentration of samples at RR = 1 indicates that the model reliably retrieves the correct cross-modal pairs.
Source-paper figure
Figure 4. Post-training results for the RONI and IOD climate indices. Line plots compare predicted values (orange) with ground-truth values (blue), while scatter plots show their correlations. RONI exhibits strong performance (R 2 = 0.969), and IOD achieves moderate performance (R 2 = 0.448).
Figure 4. Post-training results for the RONI and IOD climate indices. Line plots compare predicted values (orange) with ground-truth values (blue), while scatter plots show their correlations. RONI exhibits strong performance (R 2 = 0.969), and IOD achieves moderate performance (R 2 = 0.448).
Source-paper figure
3.3) on ERA5 data (see Sec. 3.1. The linear decoder is trained on data from 1980 -2000 and tested on data from 2001 -2015 with 2016 -2025 as the validation set. Figure 4 shows results for two representative indices from the posttraining experiment, Relative Oceanic Niño Index (RONI) and Indian Ocean Dipole (IOD), achieving R 2 scores of of 0.969 and 0.448 and a correlation of 0.98 and 0.693, respectively. RONI performs best among all the indices because it is primarily governed by oceanic variability and ocean-atmosphere coupling processes that our foundation model is designed to encode during pretraining (for additional discussion and results for all the indices, refer to Sec. C.2 in the Appendix).
3.3) on ERA5 data (see Sec. 3.1. The linear decoder is trained on data from 1980 -2000 and tested on data from 2001 -2015 with 2016 -2025 as the validation set. Figure 4 shows results for two representative indices from the posttraining experiment, Relative Oceanic Niño Index (RONI) and Indian Ocean Dipole (IOD), achieving R 2 scores of of 0.969 and 0.448 and a correlation of 0.98 and 0.693, respectively. RONI performs best among all the indices because it is primarily governed by oceanic variability and ocean-atmosphere coupling processes that our foundation model is designed to encode during pretraining (for additional discussion and results for all the indices, refer to Sec. C.2 in the Appendix).
Source-paper figure
Figure 5 highlights two indices that perform most poorly in our post-training evaluation: the real-time multivariate Madden-Julian Oscillation (MJO) components, RMM1 and RMM2. This outcome can be attributed to two main factors. First, post-training is conducted at a monthly temporal resolution, which is too coarse to resolve MJO's intraseasonal
Figure 5 highlights two indices that perform most poorly in our post-training evaluation: the real-time multivariate Madden-Julian Oscillation (MJO) components, RMM1 and RMM2. This outcome can be attributed to two main factors. First, post-training is conducted at a monthly temporal resolution, which is too coarse to resolve MJO's intraseasonal
Source-paper figure
Figure 5. Post-training results for the Real-time multivariate MJO components 1 and 2 (RMM1 and RMM2). Line plots compare predicted values (orange) with ground-truth values (blue), while scatter plots show their correlations.
Figure 5. Post-training results for the Real-time multivariate MJO components 1 and 2 (RMM1 and RMM2). Line plots compare predicted values (orange) with ground-truth values (blue), while scatter plots show their correlations.
PDFDIGEST AI

Struggling to understand complex research papers?

Upload any PDF and get instant AI-powered explanations, summaries, and visual breakdowns. Turn dense academic writing into clear, actionable insights.

Upload a Paper

Frequently Asked Questions

To complement the contrastive objective, we incorporate a lightweight auxiliary reconstruction loss on the atmospheric modality, where the input field is reconstructed from its latent representation. The reconstruction term is weighted by α ∈ [0, 0.1] to provide a stabilizing signal without.

In domains such as natural language processing and computer vision, these models have shifted the focus from task-specific solutions to unified architectures that capture underlying structure in complex systems . The task is a regression problem that is optimized using Huber loss.

This leads to an atmosphere-centric view of the Earth system, in which slowly evolving components such as the ocean are under observed and weakly constrained, especially at the timescales relevant to their variability. Our overarching goal is to develop a foundation model.

Observational datasets, while physically grounded, are limited in temporal extent (typically < 50 years) and therefore underrepresent slower components such as the ocean, sea ice, land ice, and land surface. This metric is well-suited to our setting, as it accounts for the.

The learned latent space is designed to support transfer to downstream tasks such as climate risk assessment, extreme event prediction, and climate attribution; we leave a systematic evaluation of these capabilities to future work. NIVA addresses this limitation by operating at lower.

This paper introduces NIVA, a new AI model that helps improve weather and climate predictions by understanding how different parts of the Earth system interact. It focuses on the ocean and atmosphere to make better forecasts beyond two weeks.

Related Research

Research

Compression and Localization in Reinforcement Learning for ATARI Games

This paper discusses a new approach to making reinforcement learning models smaller and more efficient, particularly for playing ATARI games.

10 min read