## Summary
Replace the three unchecked scanner conversions in
`crates/oxide/src/scanner/mod.rs` with checked conversions that omit
only extracted slices that are not valid UTF-8. Apply the check in the
shared `extract` pipeline so initial scans, incremental `scan_content`
calls, and CSS-variable extraction cannot insert invalid strings into
scanner state, and apply the same policy to both branches of
`get_candidates_with_positions` while preserving byte offsets and the
legacy `-[]` restoration. Keep the extractor's byte-oriented CSS
identifier classification unchanged: accepting non-ASCII bytes during
extraction is useful for valid multibyte code points, while the
conversion boundary is the authoritative place to enforce the `String`
contract.
The scanner currently converts extracted byte slices with unchecked
UTF-8 constructors at the shared extraction boundary and both
candidate-with-position branches. A source file containing a stray
continuation byte can therefore produce an invalid `String`, violating
Rust's string invariant and allowing corrupted candidates to persist in
a long-lived scanner. The thread provides a deterministic reproduction
using invalid bytes inside an arbitrary value, so the problem no longer
depends on reproducing the originally reported Turbopack race. Valid
candidates found alongside malformed byte sequences must continue to be
returned normally.
Fixes#20368
## Test plan
Not applicable to this change.
AI was used for assistance.
---------
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Robin Malfait <malfait.robin@gmail.com>