How Error-Mitigated Quantum Computers Are Reaching Beyond State-of-the-Art Classical Simulation

by Tasneem Watad PhD

Key takeaways

  • We studied the dynamics of a Floquet quantum magnet on heavy-hex lattices containing up to 74 spins using an IBM Heron R3 processor.
  • QESEM, Qedma’s Error Suppression and Error Mitigation software, enabled high-precision quantum simulations at intermediate physical times, a regime in which the leading classical methods we investigated, including tensor-network and Sparse Pauli Path simulations, could not capture the dynamics because of prohibitive memory and runtime requirements, despite extensive computational resources, including NVIDIA H100 GPUs and over 500,000 CPU-hours across ~13,000 nodes of the Fugaku supercomputer.
  • Our results address a physically important question beyond the reliable reach of today’s leading classical methods, providing evidence that long-lived prethermal oscillations of the studied Floquet magnet persist in large-scale systems, contrary to the intuitive expectation that they would gradually disappear as a finite-size effect.
  • The reliability of these findings is supported by multiple independent validation tests, including cross-platform validation on Quantinuum trapped-ion hardware (H2 and Helios) and unbiased error mitigation backed by rigorous theoretical accuracy guarantees, among others.
  • These results illustrate how advanced quantum error mitigation is enabling today’s quantum processors to become quantitative scientific tools for probing many-body physics beyond the reach of classical simulation methods.

Introduction

One of the central promises of quantum computing is the ability to study quantum many-body physical systems beyond the practical reach of classical computers. However, demonstrating this in a scientifically meaningful way is challenging. It is not enough to show that a quantum processor can execute a large circuit , the experiment must probe a regime beyond reliable classical prediction while producing measurement results we can trust.

Recent advances in quantum hardware have moved us substantially toward this goal, but hardware improvements alone are not enough. Advanced quantum error mitigation is becoming an equally important ingredient, extending the range of quantum circuits that can be executed with quantitative accuracy. At Qedma, we have developed QESEM, our Quantum Error Suppression and Error Mitigation software, to turn today’s quantum processors into reliable scientific instruments capable of studying complex quantum systems.

In our latest work, carried out in collaboration with IBM, RIKEN and BlueQubit, we used QESEM on an IBM Heron R3 processor to investigate the intermediate-time dynamics of a periodically driven (Floquet) quantum magnet. The experiments employed two complementary QESEM protocols, both of which are commercially available as part of Qedma’s QESEM software: QESEM-Unbiased, which combines high-accuracy device characterization with quasi-probabilistic error mitigation to produce expectation values that are unbiased up to characterization errors, together with rigorous statistical error bars that enable quantitative validation of the results; and QESEM-Extrapolated (aka QESEM-X), a lower-overhead protocol that extends the accessible evolution time beyond that reached by QESEM-Unbiased using a carefully validated heuristic procedure.

Fig. 1: Computational cost of the quantum and classical approaches considered in this work. The shaded beige region marks the regime where the state-of-the-art classical methods investigated here (statevector, tensor networks -PEPS-BP- and Sparse Pauli Paths) become computationally impractical or lose convergence. In contrast, error-mitigated quantum experiments using the QESEM-Unbiased and QESEM-Extrapolated protocols continue to provide reliable access to the dynamics beyond this regime.

By combining large-scale error-mitigated quantum experiments with extensive classical benchmarking, as demonstrated in Fig. 1, we identified a regime in which multiple state-of-the-art classical simulation methods cease to provide controlled predictions, while the error-mitigated quantum measurements of the quantum processor continue to probe the quantum system behavior reliably.

Why Floquet Systems?

One of the biggest challenges in demonstrating the power of quantum computing is identifying a problem that is simultaneously scientifically interesting, experimentally accessible, and genuinely difficult for classical computers. Periodically driven interacting quantum many-body systems provide exactly such a setting.

Unlike static systems, Floquet systems are repeatedly driven by a sequence of quantum operations. This periodic driving gives rise to rich non-equilibrium physics that has no equilibrium analogue, including long-lived prethermal states, time crystals, engineered topological phases, and artificial gauge fields. Because of these unique properties, Floquet systems have become a major research direction in condensed matter physics and quantum simulation.

Beyond their intrinsic scientific interest, Floquet systems provide an ideal benchmark for quantum computing. The Floquet mixed-field Ising model studied here rapidly generates both state entanglement and operator entanglement (operator scrambling), causing quantum information to spread quickly throughout the system. These are precisely the features that make accurate classical simulation increasingly difficult at later stages of the dynamics.

The Floquet Circuit

Fig. 2 illustrates a single Floquet cycle for a smaller 12-qubit heavy-hex patch; the full experiment follows the same structure on systems containing up to 74 qubits.

In this case, each Floquet cycle consists of three subcycles. In each subcycle, rotations about the xx– and z-axes, are applied in parallel on all qubits, followed by a layer of entangling Rzz gates acting on one of three complementary edge sets of the heavy-hex lattice. Together, the three subcycles ensure that every interaction in the lattice is applied once per Floquet cycle. Repeating this sequence for many Floquet cycles produces increasingly complex many-body dynamics while remaining compatible with the native connectivity of IBM’s heavy-hex processors.

Quantum Experiment and Classical Benchmarks

We implemented the Floquet circuit described above on an IBM Heron quantum processor and measured the magnetization of the quantum spins throughout the evolution. The experiments were performed on heavy-hex ‘ladders’ containing up to 74 qubits (the 51-qubit and 74-qubit patches are illustrated below), using the two protocols described above: QESEM-Unbiased and QESEM-Extrapolated.

Fig. 3: The 51-qubit (left) and 74-qubit (right) heavy-hex patches of the IBM Heron R3 quantum processor used for the quantum simulation experiments.

To benchmark the quantum results, we compared them with two of today’s leading classical simulation methods: Tensor Networks (PEPS-BP) and Sparse Pauli Paths (SPP). Although both are state-of-the-art approaches, they tackle the simulation from fundamentally different perspectives. PEPS-BP works in the Schrödinger picture, directly evolving the many-body quantum state as the Floquet circuit is applied layer by layer. In contrast, SPP works in the Heisenberg picture, propagating the measured observable, in our case the magnetization, backward through the circuit, where it expands into an increasingly large sum of Pauli operators.

Both methods rely on controlled approximations to remain computationally tractable. In PEPS-BP, the approximation is governed primarily by the tensor-network bond dimension, which limits how much entanglement can be represented. In SPP, only the most significant Pauli terms are retained according to their support size, WW, and coefficient magnitude, ε0\varepsilon_0 . Decreasing ε0\varepsilon_0, increasing WW and increasing the bond dimension of the PEPS generally improves the accuracy, but at rapidly growing computational cost.

As we show next, both classical approaches accurately reproduce the early stages of the dynamics. Beyond this regime, their convergence deteriorates as entanglement and operator complexity continue to grow. Besides PEPS-BP and SPP, we also investigated several additional classical simulation approaches, including PEPO-BP and a heuristic rescaling method based on one-dimensional matrix product state (MPS) simulations. These methods likewise fail to capture the dynamics at the intermediate-time regime, and therefore, we focus the discussion here on PEPS-BP and SPP.

Quantum & Classical Results

The figure below compares the error-mitigated quantum experiments with the PEPS-BP and SPP simulations.

The complete quantum experiment consumed approximately 16.4 QPU hours, in addition to classical post-processing. PEPS-BP simulations were carried out up to a bond dimension of 700, which, to the best of our knowledge, exceeds the bond dimensions typically reported in the literature for comparable two-dimensional simulations. A bond dimension of 512 represented the practical memory limit on a single NVIDIA H100 GPU; and reaching a bond dimension of 700 required approximately two weeks on a CPU cluster (roughly 8,000 effective vCPU-hours). The largest SPP simulation presented in this figure utilized 65,536 Fugaku cores (distributed over 12888 nodes) for approximately 8 hours, amounting to over 540,000 Fugaku core-hours while tracking nearly 101210^{12} Pauli strings. 

Remarkably, unlike the classical simulations, the same quantum experimental data can subsequently be used to evaluate many additional observables, such as correlation functions, without rerunning the experiment, providing further insight into the Floquet dynamics. Obtaining each of these observables classically would generally require additional simulation time beyond the resource estimates quoted above.

Fig. 4: Magnetization as a function of Floquet cycle measured on the IBM Heron R3 processor using the QESEM-Unbiased and QESEM-Extrapolated protocols, together with results from leading classical simulation methods. The dynamics naturally divide into two regimes: an initial transient followed by a long-lived prethermal regime characterized by slow relaxation and persistent oscillations, before the system eventually heats toward a featureless infinite-temperature state. The data shown here capture the transient and prethermal regimes but not the eventual heating. During the transient and the first several Floquet cycles of the prethermal regime, the classical simulations agree well with one another and with the error-mitigated quantum results. Beyond approximately twelve Floquet cycles, however, the classical methods lose controlled convergence and begin to diverge both from one another and from the quantum experiments. Over the entire range where both are available, QESEM-Unbiased and QESEM-Extrapolated remain in close agreement, after which QESEM-Extrapolated extends the evolution further and reveals persistent oscillations that are not well captured by the classical methods.

Since the experiments probe a regime where controlled classical predictions are no longer available, establishing their reliability is just as important as the physics itself. We therefore performed multiple independent validation tests, including unbiased error mitigation with rigorous theoretical guarantees, agreement between the QESEM-Unbiased and QESEM-Extrapolated protocols, agreement with classical simulations where they remain converged, validation of the hardware noise model, and cross-platform verification at selected Floquet cycles on the Quantinuum System Model H2 and Quantinuum Helios trapped-ion processors. A detailed discussion of these validation procedures can be found in our paper. 

Main findings and their significance

The central finding of this work is that the long-lived oscillatory behavior persists well beyond the regime where today’s leading classical simulation methods provide reliable predictions.

An important question is whether the persistence of the oscillations could already have been inferred from exact simulations of smaller systems. As the animation below illustrates, the answer is no.

Finite-size scaling based only on the classically accessible system sizes leaves a broad confidence interval, making it impossible to determine whether the oscillation amplitude ultimately vanishes or remains finite in the thermodynamic limit of heavy-hex ladders. Once the large-scale error-mitigated quantum experiments, including the 74-qubit measurements, are incorporated, the confidence interval narrows substantially. The combined data provides strong evidence that the oscillatory response persists as the system size increases.

Fig. 5. Left: Intermediate-time magnetization for 28-, 35-, 51-, and 74-qubit systems. The 28- and 35-qubit results are obtained from exact classical simulations, while the 51- and 74-qubit results come from large-scale error-mitigated quantum experiments. Although the oscillation amplitude decreases with increasing system size, the oscillatory response remains clearly visible in the larger systems. Right: Extrapolating the oscillation amplitudes corresponding to only the small, classically accessible systems does not determine whether the oscillations survive at larger scales. Incorporating he 51- and 74-qubit error-mitigated quantum experiments provides strong evidence that they do.

Why do the classical methods fail?

Although PEPS-BP and SPP approach the simulation from fundamentally different perspectives, the Schrödinger and Heisenberg pictures, respectively, both are ultimately strained by the rapid growth of state and operator entanglement, which drives the computational complexity from complementary directions. The natural question is therefore how much further these methods would need to be pushed before they converge.

For SPP, the answer appears to be: substantially further. Even after tracking nearly 101210^{12} Pauli strings on the Fugaku supercomputer and consuming more than 500,000500,000 Fugaku core-hours, the simulations remain far from convergence and fail to reproduce the intermediate-time oscillatory behavior. Importantly, this should not be interpreted as a limitation of SPP itself, but rather as a consequence of the exceptional complexity of the Floquet dynamics studied here. This conclusion is reinforced by smaller systems, where exact statevector calculations are still possible. There, truncated SPP simulations systematically shift, suppress, or overestimate the oscillation amplitude, and recover the correct dynamics only when they retain an enormous fraction of the full Pauli basis, as illustrated in the figure below. Extrapolating this behavior to the 51-qubit circuit suggests that convergence would require retaining on the order of 1030 Pauli strings, far beyond what is computationally feasible.

Fig. 6: SPP simulations for a 21-qubit system. Convergence requires retaining a large fraction of the Pauli basis, illustrating the rapid growth of operator complexity even in relatively small systems.

The tensor-network simulations tell a similar story. Although PEPS-BP represents one of the leading approaches for simulating two-dimensional quantum systems, the rapid growth of state entanglement eventually drives the simulation beyond its controlled convergence regime. Conservative scaling estimates based on the converged portion of the early-time dynamics suggest that accurately reaching, for example, 25 Floquet cycles would require months of runtime on roughly 10,000 servers, each equipped with eight NVIDIA H100 GPUs.

Looking Ahead

The competition between quantum and classical computation drives progress in quantum simulation, and further research may extend the classical frontier presented here. To encourage further progress, we are making these benchmark circuits available through the Quantum Advantage Tracker, inviting the community to test their best classical methods against these benchmarks. For condensed matter physics, these are exciting times. This work marks not only a computational milestone but also advances our understanding of non-equilibrium quantum matter . We are beginning to reach problems where today’s quantum processors, combined with rigorous error mitigation, can provide scientifically useful answers beyond the reach of controlled classical simulation. QESEM was developed precisely for this purpose: to enable quantum computers to produce accurate and trustworthy scientific results from today’s noisy processors to future fault-tolerant quantum computers.

If you are interested in the technical details, validation procedures, experimental methodology, and complete benchmarking study, the full preprint is available here.

Interested in what we do?

More To Explore