## Summary
This PR fixes a panic that occurs when the Ruby or Vue preprocessors
encounter files with invalid UTF-8 bytes.
**The issue:**
- `ruby.rs:37` and `vue.rs:18` used
`std::str::from_utf8(content).unwrap()`
- This panics when processing files containing invalid UTF-8 bytes
**Error message:**
```
thread panicked at crates/oxide/src/extractor/pre_processors/ruby.rs:37:59:
called `Result::unwrap()` on an `Err` value: Utf8Error { valid_up_to: 45, error_len: Some(1) }
```
**The fix:**
- Wrap UTF-8 conversion in `if let Ok(...)` to gracefully handle invalid
UTF-8
- Skip regex-based template extraction when UTF-8 conversion fails
- Allow byte-level processing to continue (in Ruby's case)
This can happen in Rails projects when:
- Binary files are inadvertently scanned
- Files contain non-UTF-8 encodings
- Files are truncated at multi-byte character boundaries during parallel
processing
## Test plan
- [x] Added `test_invalid_utf8_does_not_panic` test for Ruby
preprocessor
- [x] Added `test_valid_utf8_with_multibyte_chars` test for Ruby
preprocessor
- [x] Added `test_invalid_utf8_does_not_panic` test for Vue preprocessor
- [x] All existing tests pass (`cargo test pre_processors` - 43 tests)
---------
Co-authored-by: Robin Malfait <malfait.robin@gmail.com>
|
||
|---|---|---|
| .. | ||
| classification-macros | ||
| ignore | ||
| node | ||
| oxide | ||