Normalize trailing decimal zeros before bigint comparison
perfloop/fast_float · INEFFICIENT ALGORITHM
https://perfloop.ai/t/oss/case_xwm9h4yejp
Verdict
VERIFIED · settled 2026-08-15
What happened: The paired measurements met the required improvement.
Hypothesis
`digit_comp` derives `sci_exp` from the original parsed coefficient at digit_comparison.h:437-438, then `parse_mantissa` reconstructs every retained digit into a bigint (255-337) and line 444 derives the decimal scale as `sci_exp + 1 - digits`. If a coefficient ends in z zero digits, parsing only the prefix through its last nonzero digit reduces both the bigint coefficient and `digits` by z; line 444 increases the scale by z, preserving the exact decimal value. The current path instead builds zero limbs up to `binary_format<double>::max_digits()` (769) and can enter `negative_digit_comp` with that unnormalized representation.
The indexed incoming-call check found `from_chars_advanced` calling this target at parse_number.h:286, where the caller documents the `am.power2 < 0` route as very uncommon. The cadence is therefore once per unresolved long conversion, not per ordinary parse. A local `ff_trim_zero_diff` C++ differential generated 37 real `too_many_digits` cases for which `compute_float(m)` differed from `compute_float(m + 1)`; the current target and a version with only logical terminal-zero span trimming returned identical adjusted mantissas. A separate direct-target calibration used a 735-character coefficient with a 30-digit nonzero core, 700 terminal zeros, and a compensated exponent: across 50,000 calls it reported about 80-81 ms for the current target versus 19 ms for the trimmed variant. Those are source-local synthetic checks, not evidence of production input frequency or end-to-end latency.
A case session should exercise the public `from_chars` entry with actual unresolved long inputs while sweeping nonzero-core length, zero-suffix length (including 0, below/above the crossover, 769+, and much longer), integer versus fraction suffix placement, and explicit exponents. It should differentially verify result bits, pointer, and error code, measure target and end-to-end CPU/latency, and use branch/input telemetry to establish whether long zero-padded fields occur often enough to matter. This is local normalization after the current handoff and does not replace the already-published parse/reparse-span fix.
Change to test: Inside `digit_comp`, add a logical trailing-zero normalization for the significant-digit spans before calling `parse_mantissa`: when the final fraction or integer digit is zero, find the last nonzero digit by scanning fraction then integer backward, and pass shortened logical ends to the mantissa parser only past a measured crossover. Keep `sci_exp` from the original parsed number so the existing `sci_exp + 1 - digits` calculation absorbs the removed powers of ten. Retain the current spans for no suffix and all-zero inputs, and preserve the current truncation and exact-rounding behavior for any nonzero suffix.
Where it lives
perfloop/fast_float · include/fast_float/parse_number.h
Evidence
public fast_float::from_chars<double> unresolved long decimal inputs with 20-, 30-, and 120-digit nonzero cores and 64, 700, 769, or 4096 terminal zeroes in integer, fraction, and decimal layouts · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns_per_parse |
3157 |
1925 |
−39.1% (−1233) |
−1271 to −1220 |
< 0 |
PASSED |
public fast_float::from_chars<double> unresolved decimal inputs with 120-, 720-, and 769-digit nonzero cores and 16-63 terminal zeroes in integer, fraction, and decimal layouts · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns_per_medium_suffix_parse |
1979 |
1930 |
−2.5% (−49.38) |
−56.85 to −31.25 |
< 0 |
PASSED |
public fast_float::from_chars<double> unresolved decimal inputs with 20-, 30-, 120-, and 769-digit nonzero cores and 0-15 terminal zeroes in integer, fraction, and decimal layouts · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns_per_pre_cutoff_parse |
1196 |
1191 |
−0.4% (−5.331) |
−14.13 to +2.589 |
≤ 12 |
PASSED |
ordinary public fast_float::from_chars<double> decimal inputs that do not reach the digit-comparison fallback · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns_per_ordinary_parse |
10.79 |
10.5 |
−2.9% (−0.3165) |
−0.418 to −0.207 |
≤ 0 |
PASSED |
Checks: 2 of 2 passed. Verification: no defect found.
Timeline
2026-07-31· Case opened2026-08-04· PR opened