Skip to content

fix: preserve byte widths when skipping invalid UTF-8 in Substring - #1007

Open
jakezwang wants to merge 1 commit into
samber:masterfrom
jakezwang:fix/substring-invalid-utf8-offset
Open

jakezwang wants to merge 1 commit into
samber:masterfrom
jakezwang:fix/substring-invalid-utf8-offset

Conversation

@jakezwang

Copy link
Copy Markdown

Substring("\xffa", 1, 1) panics, and skipping an invalid byte in longer strings can discard valid trailing characters. The positive-offset path uses utf8.RuneLen on the replacement rune, which reports three bytes even when range consumed one invalid byte.

Use the decoded byte width at that position. Regressions cover the panic, incorrect truncation, consecutive invalid bytes, and a valid U+FFFD character.

Validation: the three invalid-input cases fail before the fix. Full tests pass on Go 1.18.10; full race tests pass on Go 1.26.6 and 1.27.1. Golangci-lint passes, and the existing substring benchmarks remain allocation-free. One initial Go 1.27 run missed the unrelated TestWaitFor timing tolerance; the full rerun passed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant