Evidence & scope / October 2026
Know what the
numbers describe.
Runtime, numerical accuracy and supported features answer different questions. This page records the basis of the comparisons on the MinuteSim website.
Large-model runtime comparison
The homepage presents a solid-compression case with 1,886,592 elements, a quarter model under a moving rigid hemisphere, a 500 mm tool travel and a target physical analysis duration of 0.5 s. The 590× and 92.4× values use a projected GPU runtime.
| Comparison | CPU time | GPU time | Ratio | Basis |
|---|---|---|---|---|
| 1 CPU thread | 81,303.3 s 22.6 hours | 137.8 s 2.3 minutes | 590× | Estimated CPU full duration / projected normal-clock GPU time |
| 8 CPU threads | 12,736.178 s 3.54 hours | 137.8 s 2.3 minutes | 92.4× | Recorded CPU run / projected normal-clock GPU time |
| 8 CPU threads · observed runs | 12,736.178 s | 161.519 s | 78.9× | Recorded completed CPU and GPU runs |
The one-thread CPU estimate is extrapolated from a short timing window. The 137.8 s GPU value is a supplied normal-clock projection, not a recorded complete run. The observed-run row is SUPPORTED as a runtime comparison; it is not independent validation of the result fields.
Timing evidence dates to 5–6 October 2026. Exact released-version labels are not established by this comparison. Observed CPU wall time includes the starter and engine processes; observed GPU wall time includes process startup and setup. The projected GPU figure estimates whole-run wall time, not kernel-only time; its initialization component was not separately revalidated. Startup accounting within the supplied one-thread extrapolation has not been independently confirmed.
Conditions that matter
- CPU comparator: Solver B, using the indicated thread allocation on AMD EPYC 9274F-class hardware. GPU: one NVIDIA L40 on a separate EPYC 9274F-class server.
- GPU calculation uses FP32; the CPU comparator uses its extended-single configuration. These are not identical arithmetic implementations.
- The recorded GPU run reaches approximately 0.500003 s in 119,821 explicit time increments. The recorded 8-thread CPU run uses 204,455 cycles. The website’s increment count is specific to MinuteSim.
- Time-step selection and output settings differ. GPU timing excludes animation output and between-run cooling. The observed GPU run includes some clock throttling.
- Final-field equivalence has not been established for this timing case. Model size alone does not determine relative runtime.
The homepage gallery contains separate capability demonstrations. Its forming or indentation images are not results from this 1,886,592-element timing run. Historical timings in the public documentation use other configurations and are not points on this comparison.
Runtime scenarios across model sizes
The two graphs reproduce the v11 presentation's supplied timing scenarios for sheet forming, mixed-element bending and solid compression. Each value is CPU time divided by GPU time. The panels use identical logarithmic axes and the same GPU series, with either 8 CPU threads or 1 CPU thread.
| Problem family | 8 CPU threads / GPU | 1 CPU thread / GPU |
|---|---|---|
| Sheet forming | 132× | 517× |
| Mixed-element bending | 46× | 314× |
| Solid compression | 92.4× | 590× |
These rounded labels refer to the largest plotted model in each family, not a universal or series-wide maximum. Values combine recorded timings and estimated full-run times. The 1,886,592-element compression endpoint uses the projected GPU time of 137.8 s; see the recorded full-run comparison and projection basis above.
CPU comparator: Solver B, with its extended-single-precision configuration, on EPYC 9274F-class hardware. GPU: one NVIDIA L40, FP32. Server allocation and output settings vary by case; the large compression comparison uses separate servers. Time increments differ, and GPU animation output was disabled for the timing runs. Final-field equivalence is not established by this timing comparison.
The bending timing family uses prescribed motion without contact. The contact-bending gallery and the 64,800-element AI scenario use a different configuration. Element counts across different problem families do not isolate a scaling effect. The scenario evidence was compiled on 5–6 October 2026; v11 names the presentation revision, not a solver release.
1,000 simulations for AI training
The interactive scenario multiplies a measured single-run wall time by the number of simulations. At 1,000 cases:
| Model | CPU / run | GPU / run | CPU / 1,000 | GPU / 1,000 |
|---|---|---|---|---|
| Sheet forming 19,881 shells | 6,839.55 s | 66.676 s | 79.2 days | 18.5 hours |
| Contact bending 64,800 elements | 1,429.374 s | 226.817 s | 16.5 days | 2.6 days |
CPU: Solver B, 8 threads on an AMD EPYC 9274F. GPU: one NVIDIA L40 on the same server. The GPU uses single precision and the CPU comparator its extended-single configuration. These smaller cases are separate from the large solid-compression scenario.
The calculation assumes sequential execution and constant single-run time. It excludes preprocessing, orchestration, cooling delays, dataset preparation and AI model training. These are not measured 1,000-run batches. Multiple concurrent CPU jobs, model variations and data-quality checks change the end-to-end comparison.
MeshGraphNets, Appendix A.1, provides one precedent: 1,000 training trajectories per dataset, with separate validation and test sets. This motivates an illustrative data scale; it is not a universal minimum or a guarantee of stable learning. Automated dataset generation and training integration are not established by the timings.
GPU hardware and physics data
The selected historical curves illustrate the opportunity from GPU hardware development. Their absolute linear axes show conventional FP32 TFLOPS and memory bandwidth in GB/s. Hardware specifications are not measurements of MinuteSim throughput and do not predict its speed on a newer GPU.
- FP32: selected desktop GPU published or derived peak specifications and full-chip CPU SIMD estimates at base clocks. Tensor and sparse-operation rates are excluded. GPU and CPU clock conventions are not matched.
- Memory bandwidth: selected data-center GPUs and server CPU sockets. CPU values describe the whole socket's theoretical bandwidth, not the bandwidth observed by one solver thread.
- L40: an independent reference at 90.5 TFLOPS and 864 GB/s. The 2023 marker denotes availability of L40-equipped OVX systems; the GPU was announced in 2022.
- Scope: selected products through 2025, not a complete product survey or a comparison at equal power, price or clock. Product classes and instruction mixes differ.
| Metric | GPU endpoint | CPU endpoint |
|---|---|---|
| Conventional FP32 | RTX 5090 (2025) 104.9 TFLOPS | Ryzen 9 9950X (2024) 4.4 TFLOPS |
| Memory bandwidth | B300 (2025) up to 8,000 GB/s | Xeon 6980P (2024) 844.8 GB/s |
The B300 endpoint is configuration-specific; NVIDIA documents also list 7.7 TB/s for an HGX variant. The Xeon endpoint is calculated as 12 channels × 8 bytes × 8,800 MT/s. Growth labels divide each series by its own first value: compute starts in 2008 for both series; bandwidth starts in 2008 for GPUs and 2007 for CPUs. They are not matched-period application performance growth.
Sources and attribution
Figures redrawn from selected historical data, with absolute-unit axes and an added L40 reference. Website images are crops of the v11 presentation. Sources: Karl Rupp, CPU, GPU and MIC Hardware Characteristics over Time; Epoch AI, Data on machine learning hardware (including its underlying hardware references); and selected vendor specifications. Rupp and Epoch AI data are used under CC BY 4.0.
- NVIDIA L40 specifications.
- NVIDIA HGX reference architecture and Blackwell Ultra datasheet — bandwidth varies by configuration.
- Intel Xeon 6980P specifications — memory channels and transfer rate.
Numerical evidence, case by case
Published comparisons provide evidence for their stated configurations. They do not validate every current feature or the large-model runtime scenario.
| Study case | Reference | Reported comparison | Scope |
|---|---|---|---|
| Nakajima forming | Abaqus/Explicit S4 | 2.95% mean von Mises stress difference over 94% of specimen elements; 2.08% maximum thickness difference | 10,000 shells, 80 mm stroke, FP64; shell study |
| Rounded-flat-punch normal contact | Closed-form solution | 1.69% normal-force difference | Reported coarse Tet4 configuration; solid study |
CPU/GPU precision agreement within the same solver is a self-consistency check, distinct from comparison with independent numerical or analytical references. Consult the original papers for complete methods and results, and the capability scope below for current boundaries.
What the capability list means
The homepage describes current development capabilities. It does not promise that every combination is available in the documented public beta. Confirm the required element, material, contact and constraint combination as part of an evaluation.
- Elements: hexahedral, tetrahedral, wedge, shell and beam configurations. Qualification differs by formulation and configuration; a topology being executable does not establish general accuracy or stability.
- Materials: the listed elastic, plastic, rate-dependent, anisotropic and rubber options apply to supported element/material combinations. They are not a claim of a complete commercial material library.
- Shell erosion: plastic-strain-threshold erosion with Quad4 shells is supported in selected MAT_003 and supplied MAT_024 configurations. Support remains configuration-specific. It does not establish validated fracture prediction or general erosion-keyword support.
- MPC: supported isolated, global translational linear constraints.
- RBE2: supported solid-member nodal rigid connections. Precision and small-displacement limitations apply.
- RBE3: supported global translational weighted interpolation. General local/rotational or shared-degree-of-freedom cases are not implied.
- Contact and walls: support is configuration-specific. A spherical rigid-wall demonstration is not evidence for all deformable contact.
- Adaptive shell refinement: EXPERIMENTAL. The S-rail result illustrates a configured run; no universal forming accuracy or refinement guarantee is claimed.
- Input and output: Supports a selected subset of LS-DYNA-style keyword input syntax. Results are viewable in ParaView. This is not a claim of full keyword coverage or solver equivalence.
In this website, SUPPORTED means implemented and reachable through documented input, within the stated scope; VALIDATED means a supported configuration has documented numerical comparison against an independent reference; EXPERIMENTAL identifies evaluation-stage capabilities or illustrative timing assumptions; PLANNED describes an application direction. Contact MinuteSim to discuss release-specific availability for an evaluation.
Reading the simulation gallery
Images and videos show actual MinuteSim result fields. Public derivatives remove development annotations and device logs while retaining geometry, deformation and the original color scales. Playback duration is not solver runtime.
- S-rail: effective plastic strain at the shell top fibre. Experimental adaptive refinement. The hero hides mesh lines; the animation shows the evolving mesh.
- Sheet forming: effective plastic strain in a configured forming run. The clip is an illustration, not a timing measurement or a substitute for the published validation case.
- Taylor impact: equivalent plastic strain, quarter-model view. The displayed scale runs from 0 to 2.5; larger values use the top color.
- Mixed-element bending: equivalent plastic strain in a contact configuration. It is distinct from no-contact bending performance studies.
- Tension and necking: one-layer and three-layer solid specimens with their shared original plastic-strain scale. This is not a validated fracture prediction.
- Solid tensile necking: oblique view of 7,200 hexahedral elements, three layers through the thickness, with a Swift hardening curve. Effective plastic strain uses the original 0–0.70 scale. This is a SUPPORTED configured example without an independent reference comparison.
- Shell tensile erosion: 2,400 Quad4 shells, refined before the run, with the original shell-top effective-plastic-strain scale of 0–0.50. The supplied MAT_024 configuration uses a Swift hardening curve and a failure threshold of 0.45. This is a SUPPORTED configured element-deletion example; it does not establish calibrated fracture accuracy or adaptive refinement.
- Spherical indentation: shell-plate displacement under a spherical rigid wall. A capability demonstration, not general contact validation.
Public references
- Applied Sciences 16(12), 5826 (2026) — GPU-resident shell simulation study.
- Journal of Manufacturing and Materials Processing 10(6), 197 (2026) — GPU-resident large-deformation finite-element study.
- Pfaff et al., Learning Mesh-Based Simulation with Graph Networks (ICLR 2021) — research context for simulation-based training data.
Publication links lead to the original sources. No third-party article files or benchmark decks are redistributed with this website. Hardware chart attribution is listed in the hardware section.
Typeface: Outfit, © 2021 The Outfit Project Authors. The unchanged font is distributed under the SIL Open Font License 1.1; official font distribution.