The CMIP 2026 Community Workshop in Kyoto brought together a wide range of discussions on model development, evaluation, and future priorities for the CMIP community. One session combined paleoclimate-focused presentations with broader model evaluation talks, spanning topics from paleo model–data comparison to rapid evaluation frameworks and observation uncertainty methods. A recurring message from the paleoclimate contributions was that past climates provide test cases that we cannot obtain from the instrumental record alone. They allow us to evaluate Earth system models under boundary conditions very different from today, including much warmer worlds, colder worlds, altered ice sheets, different vegetation distributions, and reorganized ocean circulation.
Across the session, speakers examined climates ranging from the Last Glacial Maximum to the Last Interglacial, the mid-Pliocene, the mid-Holocene, abrupt-127k experiments, and glacial states such as 49 ka. Although the case studies differed in timescale and region, a common thread connected them. Paleoclimate simulations are not only useful for explaining the past. They are also highly relevant for improving confidence in future projections, identifying key feedback, and guiding model development for CMIP7.
Why paleoclimate still matters for model evaluation
The session reinforced that paleoclimate offers something uniquely valuable for model evaluation: access to climate states beyond modern variability. The instrumental period is short, and future projections inevitably move toward conditions that observations do not yet sample. Paleoclimate helps bridge that gap.
An overview talk illustrated this point directly: paleoclimate data spanning the Last Glacial Maximum (with global cooling of 5–7 °C), the mid-Pliocene (with CO₂ of 360–420 ppmv and warming of 2.5–4.0 °C), and the Early Eocene (with CO₂ exceeding 1150 ppmv) collectively provide constraints on global mean surface temperature, equilibrium climate sensitivity, and polar amplification across a wide range of forcing levels. These data are already used in IPCC assessments, but they remain underused during the model development cycle itself. The concept of a “paleo-tuning grand challenge” was presented, in which paleoclimate evaluation is embedded directly into the iterative process of model improvement, not applied only retrospectively after a model generation is complete.
This perspective was reinforced across the session. Past warm and cold climates can help constrain the plausibility of simulated sensitivity, polar amplification, cloud feedbacks, and hydrological responses. In this sense, paleoclimate is a practical tool for improving models, not simply a retrospective exercise.
The Arctic as a focal point for past–future comparison
A second theme concerned the Arctic, where multiple talks converged on how warm past climates can inform our understanding of future high-latitude change. The Last Interglacial featured prominently in these discussions because it combines evidence for strong summer Arctic warmth, reduced sea ice, and major cryosphere and ecosystem changes.
One presentation introduced the CMIP7 Fast Track abrupt-127k experimental protocol, designed as a short (100-year), computationally inexpensive simulation that abruptly imposes 127 ka orbital and greenhouse gas forcing from a piControl state. The experiment focuses on the strong insolation anomaly in the Arctic, up to 60–80 W m⁻² additional top-of-atmosphere insolation during May–June, and its impact on sea ice. Results showed that six of ten CMIP6 models already simulate a seasonally sea-ice-free Arctic under 127 ka conditions. This matters because future projections still show a large spread in the timing of an ice-free Arctic, and observations alone do not cover the expected future state. The abrupt-127k experiment therefore provides an additional benchmark for assessing whether models capture the relevant sensitivities and feedbacks related to Arctic sea ice loss.
Another talk addressed a key question about why models differ in their Arctic simulations: the role of cloud-phase parameterization. By comparing two representations of the supercooled liquid fraction (SLF) in mixed-phase clouds, the study showed that the parameterization with more supercooled liquid water at lower temperatures produced thinner baseline sea ice, larger sea-ice reductions at the LIG, and a stronger cloud greenhouse effect in early winter. The effect was large enough to matter: when combined with dynamic vegetation feedbacks, this cloud-phase representation led to a summer sea-ice-free Arctic at the LIG. Cloud-phase representation thus appears to be an important factor for both paleoclimate and future Arctic simulations.
The Arctic discussion also extended beyond sea ice itself. One presentation framed the Last Interglacial as a paleoclimate counterpart for the TIPMIP-WhatIf boreal forest experiment. During the LIG, pollen and plant macrofossil evidence shows that boreal forest extended to the Arctic coast across much of North America and Eurasia, permafrost was absent (based on speleothems in Siberian caves), the Greenland ice sheet was substantially smaller, and the Arctic was seasonally sea-ice free. Coupled model experiments with different vegetation prescriptions demonstrated that climate–vegetation feedbacks significantly amplified Greenland ice sheet retreat, with the vegetation-aware simulation producing a global mean sea level contribution of 3.0 m compared to just 0.6 m with preindustrial vegetation. This highlights that Arctic change involves cascading interactions among vegetation, albedo, sea ice, and ice sheets, interactions that are particularly relevant for assessing future Arctic risk.
Hydroclimate and circulation responses depend on mean state
Beyond the Arctic, several talks emphasized that the hydroclimate response to forcing depends strongly on the underlying climate state.
One presentation addressed the long-standing paradox in future monsoon projections, in which South Asian summer monsoon rainfall increases while circulation weakens. By examining PMIP simulations across the mid-Pliocene, Last Interglacial, and mid-Holocene using six model groups and 42 simulations, the study showed that consistent monsoon responses emerge across these warm intervals: an overall increase in rainfall, weakening over the Bay of Bengal, and strengthening over the northern Arabian Sea. The thermodynamic component follows the “wet gets wetter” paradigm and scales with global mean surface temperature (r = 0.92), while the dynamic component is driven by regional meridional temperature contrast. Physics-based regression models constructed from past climate information reproduce future monsoon projections with spatial pattern correlations of 0.8 for circulation and 0.7 for rainfall. This demonstrates that paleoclimate states can directly inform the physical interpretation of future monsoon change.
A related study using ACCESS-ESM1.5 examined ENSO and tropical hydroclimate variability across multiple paleoclimate states (pre-industrial, 8.2 ka, 49 ka, and the Last Interglacial), finding that ENSO characteristics and its response to AMOC weakening vary substantially with background climate, underscoring that internal variability is not an invariant property of the climate system but depends on mean state.
Taken together, these studies highlighted a broader lesson: if we want confidence in future regional hydroclimate projections, we need models that behave credibly across multiple climate states, not only under present-day conditions.
New tracers, new diagnostics, and robust evaluation
The session also considered the use of tracers and proxy-relevant variables to deepen model evaluation. One presentation showed results from three isotope-enabled fully coupled models (MPI-ESM-wiso, AWI-ESM-wiso, and MIROC6-iso) run under the abrupt-127k protocol. These simulations showed that while Arctic summer temperatures increase under 127 ka forcing, the δ18O response in precipitation also increases, though the relationship between temperature and isotopic content varies across models and is particularly complex near sea-ice boundaries. The isotope–temperature gradient is notably larger at sea-ice area boundaries, where changes in moisture source, transport, and phase changes complicate the standard interpretation.
This type of evaluation is especially valuable in polar regions, where proxy interpretation often depends on transport, seasonality, source conditions, and phase changes. Rather than evaluating only temperature or precipitation, isotope diagnostics help track the pathways and transformations of water through the coupled climate system. Isotope-enabled modeling in this way strengthens the dialogue between PMIP-style paleoclimate experiments and the broader CMIP community, creating a more process-aware bridge between model output and paleodata.
A related challenge, relevant to both modern and paleoclimate evaluation, is how to account for observational uncertainty when assessing model performance. The Observation Range Adjusted (ORA) method was presented as one approach. Model evaluation results can depend strongly on which observational dataset is used, and ORA addresses this by comparing model output against multiple observational estimates simultaneously: when a model falls within the range of observations, it is considered indistinguishable from observations and assigned no error; when it falls outside, the error is measured as the distance to the nearest observation. Applied to CMIP6 precipitation and temperature using datasets such as REGEN, MSWEP, GPCC, CRU, and Berkeley, the results showed that model performance rankings can shift depending on the reference dataset, and that ORA provides a more robust assessment. This matters for paleoclimate evaluation as well, where proxy reconstructions carry even larger uncertainties than instrumental records. As the field moves toward more systematic model–data comparison, methods that explicitly account for observation and reconstruction uncertainty will become increasingly important.
What this means for CMIP7
For the CMIP community, the session showed that paleoclimate does not provide easy answers, but it does provide demanding and highly informative tests. It pushes models into regimes where key feedback become more visible, and where differences in cloud physics, sea ice processes, vegetation coupling, and circulation dynamics can be more clearly identified.
This is why paleoclimate deserves a more central role in CMIP7. If model development is to become more targeted, rapid, and process-based, as the “paleo-tuning grand challenge” concept envisions, then evaluation needs to extend beyond the historical period alone. The abrupt-127k experiment, designed to be straightforward to set up and cheap to run, exemplifies how paleo experiments are increasingly being designed not as isolated exercises but as part of a wider strategy for model benchmarking and process understanding. Past climates provide independent benchmarks for assessing model realism under strong forcing, different geography, and alternative equilibrium states. They can help identify which model improvements are likely to matter most for future projections.
The session made clear that this is already happening. Paleo experiments are increasingly being designed alongside broader model evaluation infrastructure, including rapid evaluation frameworks, robust observation-aware metrics, and cross-timescale diagnostics, as part of a comprehensive strategy for CMIP7 model development. Looking to past climates is, in practice, helping the community prepare for what lies ahead.