Normalize trailing decimal zeros before bigint comparison

perfloop/fast_float · INEFFICIENT ALGORITHM

https://perfloop.ai/t/oss/case_xwm9h4yejp

Verdict

VERIFIED · settled 2026-08-15

What happened: The paired measurements met the required improvement.

Hypothesis

`digit_comp` derives `sci_exp` from the original parsed coefficient at digit_comparison.h:437-438, then `parse_mantissa` reconstructs every retained digit into a bigint (255-337) and line 444 derives the decimal scale as `sci_exp + 1 - digits`. If a coefficient ends in z zero digits, parsing only the prefix through its last nonzero digit reduces both the bigint coefficient and `digits` by z; line 444 increases the scale by z, preserving the exact decimal value. The current path instead builds zero limbs up to `binary_format<double>::max_digits()` (769) and can enter `negative_digit_comp` with that unnormalized representation.

The indexed incoming-call check found `from_chars_advanced` calling this target at parse_number.h:286, where the caller documents the `am.power2 < 0` route as very uncommon. The cadence is therefore once per unresolved long conversion, not per ordinary parse. A local `ff_trim_zero_diff` C++ differential generated 37 real `too_many_digits` cases for which `compute_float(m)` differed from `compute_float(m + 1)`; the current target and a version with only logical terminal-zero span trimming returned identical adjusted mantissas. A separate direct-target calibration used a 735-character coefficient with a 30-digit nonzero core, 700 terminal zeros, and a compensated exponent: across 50,000 calls it reported about 80-81 ms for the current target versus 19 ms for the trimmed variant. Those are source-local synthetic checks, not evidence of production input frequency or end-to-end latency.

A case session should exercise the public `from_chars` entry with actual unresolved long inputs while sweeping nonzero-core length, zero-suffix length (including 0, below/above the crossover, 769+, and much longer), integer versus fraction suffix placement, and explicit exponents. It should differentially verify result bits, pointer, and error code, measure target and end-to-end CPU/latency, and use branch/input telemetry to establish whether long zero-padded fields occur often enough to matter. This is local normalization after the current handoff and does not replace the already-published parse/reparse-span fix.

Change to test: Inside `digit_comp`, add a logical trailing-zero normalization for the significant-digit spans before calling `parse_mantissa`: when the final fraction or integer digit is zero, find the last nonzero digit by scanning fraction then integer backward, and pass shortened logical ends to the mantissa parser only past a measured crossover. Keep `sci_exp` from the original parsed number so the existing `sci_exp + 1 - digits` calculation absorbs the removed powers of ten. Retain the current spans for no suffix and all-zero inputs, and preserve the current truncation and exact-rounding behavior for any nonzero suffix.

Where it lives

perfloop/fast_float · include/fast_float/parse_number.h

Evidence

public fast_float::from_chars<double> unresolved long decimal inputs with 20-, 30-, and 120-digit nonzero cores and 64, 700, 769, or 4096 terminal zeroes in integer, fraction, and decimal layouts · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns_per_parse 3157 1925 −39.1% (−1233) −1271 to −1220 < 0 PASSED

public fast_float::from_chars<double> unresolved decimal inputs with 120-, 720-, and 769-digit nonzero cores and 16-63 terminal zeroes in integer, fraction, and decimal layouts · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns_per_medium_suffix_parse 1979 1930 −2.5% (−49.38) −56.85 to −31.25 < 0 PASSED

public fast_float::from_chars<double> unresolved decimal inputs with 20-, 30-, 120-, and 769-digit nonzero cores and 0-15 terminal zeroes in integer, fraction, and decimal layouts · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns_per_pre_cutoff_parse 1196 1191 −0.4% (−5.331) −14.13 to +2.589 ≤ 12 PASSED

ordinary public fast_float::from_chars<double> decimal inputs that do not reach the digit-comparison fallback · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns_per_ordinary_parse 10.79 10.5 −2.9% (−0.3165) −0.418 to −0.207 ≤ 0 PASSED

Checks: 2 of 2 passed. Verification: no defect found.

Timeline