Compare commits

...
Sign in to create a new pull request.

23 commits

Author SHA1 Message Date
Robin Malfait
ea25394532
visualize new walker tests 2026-08-13 11:35:42 +02:00
Robin Malfait
7c89b317d6
more cases 2026-08-13 11:35:42 +02:00
Robin Malfait
6ecef9db76
fix more stuff 2026-08-13 11:35:42 +02:00
Robin Malfait
a547b31500
refactor 2026-08-13 11:35:42 +02:00
Robin Malfait
5360b0d408
refactor 2026-08-13 11:35:42 +02:00
Robin Malfait
97b7d342f8
re-add External 2026-08-13 11:35:42 +02:00
Robin Malfait
c0ab06d515
let later directory sources re-include earlier @source not exclusions
`@source not "./src"` followed by `@source "./src"` scanned nothing:
in the auto/external walker, `@source not` directives were registered as
explicit gitignore layers, but plain directory sources emit no layer at
all, so there was nothing a later directive could win with — the exclusion
applied regardless of order. (A counter-whitelist layer would be wrong
too: it would rank above the `.gitignore` files on disk and bypass them
inside the re-included directory.)

Handle `@source not` in the auto/external walker with a filter closure
instead, mirroring how the pattern walkers already resolve ordering: the
last directive that covers a path wins. When that is a `@source not` the
entry is excluded; when it is a later auto/external source the entry falls
through to the normal gitignore + default rules handling, so re-included
directories keep their regular auto source semantics. Excluded directories
can still be pruned safely, because every auto/external base is its own
walk root and stays reachable even when it is nested inside an excluded
directory.

While moving the exclusion out of the gitignore layers, directory-shaped
directives now also have to exclude their base directory itself (the
normalized `/**/*` pattern only matches the directory's contents), so the
directory is pruned and doesn't widen the generated watch globs.
2026-08-13 11:35:42 +02:00
Robin Malfait
7536feb801
add test for directory source ordering 2026-08-13 11:35:41 +02:00
Robin Malfait
367b59371f
ignore unreachable gitignore whitelists when promoting sources
Whether an explicitly listed directory is git ignored (and should bypass
the ignore rules as an external source) was decided by walking up from the
directory and letting the nearest `.gitignore` with a definitive answer
win. That's only half of git's precedence: a whitelist is unreachable when
a parent directory is excluded — git never descends into an excluded
directory, so re-include rules inside of it have no effect. With

    .gitignore              parent/
    parent/.gitignore       !child/

`@source "./parent/child"` stayed a regular auto source (the `!child/`
whitelist answered first), so the `.gitignore` files inside it kept
applying, even though git considers the whole tree ignored.

Decide exclusion the way git does instead: walk the path from the top
down, settling for every directory along the way whether it is excluded —
the first excluded directory makes everything below it ignored. Within a
single directory's decision the deepest `.gitignore` still wins, so
directories re-included by a deeper, reachable `!dir` pattern stay regular
auto sources. Both directions are now pinned by unit tests next to the
other promotion tests.
2026-08-13 11:35:41 +02:00
Robin Malfait
7de8663522
add test for unreachable gitignore whitelists 2026-08-13 11:35:41 +02:00
Robin Malfait
ae45aca295
always include concrete file sources, even default-ignored ones
`@source ".env"` scanned nothing: the pattern walker only bypasses the
default file rules when the glob pins an extension, and Rust's
`Path::extension()` returns `None` for leading-dot names like `.env`, so
the file fell through to the default rules — which ignore `.env` — and was
dropped.

A pattern without any wildcards names a concrete file and is the most
explicit source there is, so it now always bypasses the default file rules,
extension or not. Extension-pinning globs behave as before, and globs that
pin nothing (e.g. `blog/*/**/*`) still apply the default rules.
2026-08-13 11:35:41 +02:00
Robin Malfait
79b2ce7af8
add test for explicit extensionless file sources 2026-08-13 11:35:41 +02:00
Robin Malfait
d78a138ca3
exclude whole directories when a @source not glob matches them
`@source not "./src/ba*"` is normalized to base `src` + pattern `/ba*`
(the wildcard can't be hoisted into the base), but the pattern-walker's
not-directive check only glob-matched the file path itself: `/ba*` never
matches `/bar/index.html` since `*` doesn't cross `/`, and the directory
pruning only recognized the normalized `/**/*` shape. So wildcard directory
exclusions were silently ignored by pattern sources.

Follow gitignore semantics instead: a pattern that matches a directory
excludes the whole subtree. `NotRule::matches` now tests the path itself
and every ancestor directory up to the directive's base, and replaces both
the file-level check and the directory pruning check (the `/**/*` sentinel
special case falls out naturally). The auto/external walker already behaved
correctly because on-disk gitignore semantics apply there natively.
2026-08-13 11:35:41 +02:00
Robin Malfait
fb43708997
add test for wildcard directory source exclusions 2026-08-13 11:35:41 +02:00
Robin Malfait
35d2ea025d
simplify SourceEntry
- Fold `External { base }` into `Auto { base, external: bool }`. External
  sources were never a distinct kind of source: they are auto sources whose
  directory is itself ignored but was explicitly listed (set either directly
  for `node_modules`-style directories, or by the gitignore promotion walk).
  Every consumer except the walker-layer emission matched
  `Auto { base } | External { base }`; those pairings collapse to a single
  arm, and the promotion becomes setting a field instead of swapping
  variants.

- Delete the unused `SourceEntry` ⇄ `GlobEntry` `From` conversions; the
  glob accessors build `GlobEntry` values inline and nothing ever converted
  through these impls.

- Document that directory-shaped `@source not` directives are normalized to
  a `/**/*` pattern (semantically identical), which `NotRule::covers_dir`
  relies on.

- Make the folder check in `From<PublicSourceEntry>` honest: joining the
  "/"-pinned pattern onto the base discarded the base entirely (an absolute
  path replaces it), so the check accidentally tested a nonsense path like
  `/index.html` against the filesystem root. Strip the leading slash and
  document that this arm only matters when `optimize()` could not
  canonicalize the base.
2026-08-13 11:35:40 +02:00
Robin Malfait
cdcbf98744
questionable test 2026-08-13 11:35:40 +02:00
Robin Malfait
6d9d3b65e4
rewrite source scanning around one walker per source kind
Previously all `@source` directives were compiled into a single file walker:
directives became layered gitignore rules, `expand_restricted_patterns`
injected `*` / `/*` / `!/dir/` counter-rules to restrict concrete patterns,
and a global `filter_entry` re-checked files against every source. Because
all rules were global to the one walker, sources contaminated each other:
pattern-driven re-includes leaked into auto source results, and auto source
bases legitimized files that only a pattern's walk root reached. This caused
a family of bugs:

- #18870: globs through directories that are ignored from within (Laravel's
  `storage/` layout) found nothing
- a concrete file `@source` pointing behind `node_modules` or a git ignored
  directory scanned all of its siblings
- `@source "./foo/*.html"` on an ignored folder also scanned other
  extensions in `foo`
- external sources (explicitly listed ignored directories) scanned nested
  `node_modules`

Sources are now split over multiple walkers:

- One gitignore-aware walker for all auto/external roots. Default rules and
  `@source not` rules are registered as explicit in-memory gitignores in
  directive order (later directives win); external roots get a `!/**/*`
  whitelist plus a re-statement of the default rules so that e.g. nested
  `node_modules` stay ignored.

- One walker per pattern base. The static prefix of a glob is the explicit
  part: it is used as the walk root and `.gitignore` files above it never
  apply, even when the base is hidden inside an ignored directory. The
  wildcard part is not explicit: `.gitignore` files inside the subtree and
  the default-ignored directories still prune (`@source "./**/*.html"` does
  not descend into `node_modules` or a git ignored `dist/";
  `@source "./dist/**/*.html"` does). Matching files always win over
  file-level ignores, and globs that don't pin an extension re-apply the
  default extension rules. A filter closure makes the exact per-file
  decision, honoring the relative order of `@source not` directives, and
  `dir_could_contain_matches` keeps the walk pruned to directories that can
  actually contribute files.

The `@source` semantics are now documented as a spec at the top of
`scanner/mod.rs` and pinned by eleven new tests (six of which failed on the
old implementation). Four unit tests asserting the internals of the removed
`expand_restricted_patterns` mechanism were deleted; the user-visible
behavior they guarded is covered by the scanner-level tests.
2026-08-13 11:35:40 +02:00
Robin Malfait
5fcd42dd9f
don't promote re-included directories to external sources
When deciding whether an explicitly listed directory is git ignored (and
should therefore bypass the ignore rules as an "external" source), follow
git's precedence: walk up from the directory and let the nearest
`.gitignore` with a definitive answer win. A directory that a deeper
`.gitignore` re-includes via a `!dir` pattern is not ignored, even when an
ancestor `.gitignore` ignores it, so it keeps its regular auto source
detection behavior instead of bypassing the gitignore rules inside of it.
2026-08-13 11:34:19 +02:00
Robin Malfait
54e2c1ab15
render symlinks relatively 2026-08-13 11:34:19 +02:00
Robin Malfait
0308dd4adb
fix broken symlink tests 2026-08-13 11:34:19 +02:00
Robin Malfait
e0d271bb62
render broken paths (symlinks) 2026-08-13 11:34:18 +02:00
Robin Malfait
44765c4167
visualize scanned files 2026-08-13 11:34:18 +02:00
Robin Malfait
a4d3f804ce
rebase ignore crate from 0.4.24 to 0.4.33 2026-08-13 11:34:18 +02:00
19 changed files with 5908 additions and 1283 deletions

71
Cargo.lock generated
View file

@ -40,7 +40,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "531a9155a481e2ee699d4f98f43c0ca4ff8ee1bfd55c31e9e98fb29d2b176fe0"
dependencies = [
"memchr",
"regex-automata 0.4.8",
"regex-automata 0.4.18",
"serde",
]
@ -59,6 +59,17 @@ dependencies = [
"syn",
]
[[package]]
name = "console"
version = "0.16.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4fe5f465a4f6fee88fad41b85d990f84c835335e85b5d9e6e63e0d06d28cba7c"
dependencies = [
"encode_unicode",
"libc",
"windows-sys 0.61.2",
]
[[package]]
name = "convert_case"
version = "0.11.0"
@ -126,6 +137,12 @@ version = "1.8.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7fcaabb2fef8c910e7f4c7ce9f67a1283a1715879a7c230ca9d6d1ae31f16d91"
[[package]]
name = "encode_unicode"
version = "1.0.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "34aa73646ffb006b8f5147f3dc182bd4bcb190227ce861fc4a4844bf8e3cb2c0"
[[package]]
name = "errno"
version = "0.3.9"
@ -241,14 +258,14 @@ dependencies = [
[[package]]
name = "globset"
version = "0.4.17"
version = "0.4.20"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "eab69130804d941f8075cfd713bf8848a2c3b3f201a9457a11e6f87e1ab62305"
checksum = "07c34a9410465b45bd9787443bc7370f37735bad04b0f0cd57ff1a3186c98988"
dependencies = [
"aho-corasick",
"bstr",
"log",
"regex-automata 0.4.8",
"regex-automata 0.4.18",
"regex-syntax 0.8.5",
]
@ -273,7 +290,7 @@ dependencies = [
"globset",
"log",
"memchr",
"regex-automata 0.4.8",
"regex-automata 0.4.18",
"same-file",
"walkdir",
"winapi-util",
@ -281,7 +298,7 @@ dependencies = [
[[package]]
name = "ignore"
version = "0.4.24"
version = "0.4.33"
dependencies = [
"bstr",
"crossbeam-channel",
@ -290,12 +307,24 @@ dependencies = [
"globset",
"log",
"memchr",
"regex-automata 0.4.8",
"regex-automata 0.4.18",
"same-file",
"walkdir",
"winapi-util",
]
[[package]]
name = "insta"
version = "1.48.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "86f0f8fee8c926415c58d6ae43a08523a26faccb2323f5e6b644fe7dd4ef6b82"
dependencies = [
"console",
"once_cell",
"similar",
"tempfile",
]
[[package]]
name = "itertools"
version = "0.11.0"
@ -445,9 +474,9 @@ dependencies = [
[[package]]
name = "once_cell"
version = "1.19.0"
version = "1.21.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3fdb12b2476b595f9358c5161aa467c2438859caa136dec86c26fdd2efe17b92"
checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50"
[[package]]
name = "overload"
@ -517,7 +546,7 @@ checksum = "b544ef1b4eac5dc2db33ea63606ae9ffcfac26c1416a2806ae0bf5f56b201191"
dependencies = [
"aho-corasick",
"memchr",
"regex-automata 0.4.8",
"regex-automata 0.4.18",
"regex-syntax 0.8.5",
]
@ -532,9 +561,9 @@ dependencies = [
[[package]]
name = "regex-automata"
version = "0.4.8"
version = "0.4.18"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "368758f23274712b504848e9d5a6f010445cc8b87a7cdb4d7cbee666c1288da3"
checksum = "ad8553b9b26413251cbf30e620595c7a41b3887f03da04579c0e6b0d6a06b4b2"
dependencies = [
"aho-corasick",
"memchr",
@ -602,6 +631,12 @@ dependencies = [
"lazy_static",
]
[[package]]
name = "similar"
version = "2.7.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bbbb5d9659141646ae647b42fe094daf6c6192d1620870b449d9557f748b2daa"
[[package]]
name = "slab"
version = "0.4.12"
@ -646,7 +681,8 @@ dependencies = [
"dunce",
"fast-glob",
"globwalk",
"ignore 0.4.24",
"ignore 0.4.33",
"insta",
"log",
"pretty_assertions",
"rayon",
@ -832,6 +868,15 @@ dependencies = [
"windows-targets",
]
[[package]]
name = "windows-sys"
version = "0.61.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ae137229bcbd6cdf0f7b80a31df61766145077ddf49416a728b02cb3921ff3fc"
dependencies = [
"windows-link",
]
[[package]]
name = "windows-targets"
version = "0.52.6"

View file

@ -1,6 +1,6 @@
[package]
name = "ignore"
version = "0.4.24" #:version
version = "0.4.33" #:version
authors = ["Andrew Gallant <jamslam@gmail.com>"]
description = """
A fast library for efficiently matching ignore files such as `.gitignore`
@ -12,7 +12,10 @@ repository = "https://github.com/BurntSushi/ripgrep/tree/master/crates/ignore"
readme = "README.md"
keywords = ["glob", "ignore", "gitignore", "pattern", "file"]
license = "Unlicense OR MIT"
# CHANGED: Use an explicit edition instead of `edition.workspace = true` since this crate is
# vendored into the Tailwind CSS workspace.
edition = "2024"
rust-version = "1.88"
[lib]
name = "ignore"
@ -20,15 +23,17 @@ bench = false
[dependencies]
crossbeam-deque = "0.8.3"
globset = "0.4.17"
# CHANGED: Use the published globset crate instead of a path dependency.
globset = "0.4.20"
log = "0.4.20"
memchr = "2.6.3"
same-file = "1.0.6"
walkdir = "2.4.0"
# CHANGED: Added `dunce` to canonicalize paths without UNC prefixes on Windows.
dunce = "1.0.5"
[dependencies.regex-automata]
version = "0.4.0"
version = "0.4.18"
default-features = false
features = ["std", "perf", "syntax", "meta", "nfa", "hybrid", "dfa-onepass"]

View file

@ -18,9 +18,7 @@ fn main() {
let stdout_thread = std::thread::spawn(move || {
let mut stdout = std::io::BufWriter::new(std::io::stdout());
for dent in rx {
stdout
.write_all(&Vec::from_path_lossy(dent.path()))
.unwrap();
stdout.write_all(&Vec::from_path_lossy(dent.path())).unwrap();
stdout.write_all(b"\n").unwrap();
}
});

View file

@ -47,6 +47,7 @@ pub(crate) const DEFAULT_TYPES: &[(&[&str], &[&str])] = &[
(&["cml"], &["*.cml"]),
(&["coffeescript"], &["*.coffee"]),
(&["config"], &["*.cfg", "*.conf", "*.config", "*.ini"]),
(&["container"], &["*Containerfile*", "*Dockerfile*"]),
(&["coq"], &["*.v"]),
(&["cpp"], &[
"*.[ChH]", "*.cc", "*.[ch]pp", "*.[ch]xx", "*.hh", "*.inl",
@ -109,6 +110,7 @@ pub(crate) const DEFAULT_TYPES: &[(&[&str], &[&str])] = &[
(&["hbs"], &["*.hbs"]),
(&["hs"], &["*.hs", "*.lhs"]),
(&["html"], &["*.htm", "*.html", "*.ejs"]),
(&["hurl"], &["*.hurl"]),
(&["hy"], &["*.hy"]),
(&["idris"], &["*.idr", "*.lidr"]),
(&["janet"], &["*.janet"]),
@ -185,6 +187,7 @@ pub(crate) const DEFAULT_TYPES: &[(&[&str], &[&str])] = &[
(&["mint"], &["*.mint"]),
(&["mk"], &["mkfile"]),
(&["ml"], &["*.ml"]),
(&["mojo"], &["*.mojo"]),
(&["motoko"], &["*.mo"]),
(&["msbuild"], &[
"*.csproj", "*.fsproj", "*.vcxproj", "*.proj", "*.props", "*.targets",
@ -206,11 +209,12 @@ pub(crate) const DEFAULT_TYPES: &[(&[&str], &[&str])] = &[
"*.php", "*.php3", "*.php4", "*.php5", "*.php7", "*.php8",
"*.pht", "*.phtml"
]),
(&["pkgbuild"], &["PKGBUILD"]),
(&["po"], &["*.po"]),
(&["pod"], &["*.pod"]),
(&["postscript"], &["*.eps", "*.ps"]),
(&["prolog"], &["*.pl", "*.pro", "*.prolog", "*.P"]),
(&["protobuf"], &["*.proto"]),
(&["proto", "protobuf"], &["*.proto"]),
(&["ps"], &["*.cdxml", "*.ps1", "*.ps1xml", "*.psd1", "*.psm1"]),
(&["puppet"], &["*.epp", "*.erb", "*.pp", "*.rb"]),
(&["purs"], &["*.purs"]),
@ -231,6 +235,7 @@ pub(crate) const DEFAULT_TYPES: &[(&[&str], &[&str])] = &[
(&["red"], &["*.r", "*.red", "*.reds"]),
(&["rescript"], &["*.res", "*.resi"]),
(&["robot"], &["*.robot"]),
(&["rocq"], &["*.v"]),
(&["rst"], &["*.rst"]),
(&["ruby"], &[
// Idiomatic files
@ -274,6 +279,7 @@ pub(crate) const DEFAULT_TYPES: &[(&[&str], &[&str])] = &[
(&["spark"], &["*.spark"]),
(&["spec"], &["*.spec"]),
(&["sql"], &["*.sql", "*.psql"]),
(&["ssa"], &["*.ssa"]),
(&["stylus"], &["*.styl"]),
(&["sv"], &["*.v", "*.vg", "*.sv", "*.svh", "*.h"]),
(&["svelte"], &["*.svelte", "*.svelte.ts"]),
@ -359,4 +365,14 @@ mod tests {
previous_name = name;
}
}
#[test]
fn default_types_aliases_are_sorted() {
for (aliases, _) in DEFAULT_TYPES.iter() {
assert!(
aliases.is_sorted(),
"this alias list is not sorted: {aliases:?}",
);
}
}
}

File diff suppressed because it is too large Load diff

View file

@ -102,7 +102,9 @@ impl Gitignore {
///
/// Note that I/O errors are ignored. For more granular control over
/// errors, use `GitignoreBuilder`.
pub fn new<P: AsRef<Path>>(gitignore_path: P) -> (Gitignore, Option<Error>) {
pub fn new<P: AsRef<Path>>(
gitignore_path: P,
) -> (Gitignore, Option<Error>) {
let path = gitignore_path.as_ref();
let parent = path.parent().unwrap_or(Path::new("/"));
let mut builder = GitignoreBuilder::new(parent);
@ -123,6 +125,17 @@ impl Gitignore {
/// The global config file path is specified by git's `core.excludesFile`
/// config option.
///
/// # Behavior
///
/// This routine does its best to discover any global git exclude files.
/// This will try to parse out the `excludesFile` config option in your
/// global git configuration, if necessary.
///
/// The specific things this routine tries (which are subject to change
/// based on how git behaves) are:
///
///
///
/// Git's config file location is `$HOME/.gitconfig`. If `$HOME/.gitconfig`
/// does not exist or does not specify `core.excludesFile`, then
/// `$XDG_CONFIG_HOME/git/ignore` is read. If `$XDG_CONFIG_HOME` is not
@ -145,7 +158,8 @@ impl Gitignore {
num_ignores: 0,
num_whitelists: 0,
matches: None,
// CHANGED: Add a flag to have Gitignore rules that apply only to files.
// CHANGED: Add a flag to have Gitignore rules that apply only to
// files.
only_on_files: false,
}
}
@ -190,7 +204,11 @@ impl Gitignore {
/// determined by a common suffix of the directory containing this
/// gitignore) is stripped. If there is no common suffix/prefix overlap,
/// then `path` is assumed to be relative to this matcher.
pub fn matched<P: AsRef<Path>>(&self, path: P, is_dir: bool) -> Match<&Glob> {
pub fn matched<P: AsRef<Path>>(
&self,
path: P,
is_dir: bool,
) -> Match<&Glob> {
if self.is_empty() {
return Match::None;
}
@ -243,11 +261,16 @@ impl Gitignore {
}
/// Like matched, but takes a path that has already been stripped.
fn matched_stripped<P: AsRef<Path>>(&self, path: P, is_dir: bool) -> Match<&Glob> {
fn matched_stripped<P: AsRef<Path>>(
&self,
path: P,
is_dir: bool,
) -> Match<&Glob> {
if self.is_empty() {
return Match::None;
}
// CHANGED: Rules marked as only_on_files can not match against directories.
// CHANGED: Rules marked as only_on_files can not match against
// directories.
if self.only_on_files && is_dir {
return Match::None;
}
@ -270,7 +293,10 @@ impl Gitignore {
/// Strips the given path such that it's suitable for matching with this
/// gitignore matcher.
fn strip<'a, P: 'a + AsRef<Path> + ?Sized>(&'a self, path: &'a P) -> &'a Path {
fn strip<'a, P: 'a + AsRef<Path> + ?Sized>(
&'a self,
path: &'a P,
) -> &'a Path {
let mut path = path.as_ref();
// A leading ./ is completely superfluous. We also strip it from
// our gitignore root path, so we need to strip it from our candidate
@ -326,7 +352,8 @@ impl GitignoreBuilder {
globs: vec![],
case_insensitive: false,
allow_unclosed_class: true,
// CHANGED: Add a flag to have Gitignore rules that apply only to files.
// CHANGED: Add a flag to have Gitignore rules that apply only to
// files.
only_on_files: false,
}
}
@ -337,18 +364,21 @@ impl GitignoreBuilder {
pub fn build(&self) -> Result<Gitignore, Error> {
let nignore = self.globs.iter().filter(|g| !g.is_whitelist()).count();
let nwhite = self.globs.iter().filter(|g| g.is_whitelist()).count();
let set = self.builder.build().map_err(|err| Error::Glob {
glob: None,
err: err.to_string(),
})?;
let set = self
.builder
.build()
.map_err(|err| Error::Glob { glob: None, err: err.to_string() })?;
Ok(Gitignore {
set,
root: self.root.clone(),
globs: self.globs.clone(),
num_ignores: nignore as u64,
num_whitelists: nwhite as u64,
matches: Some(Arc::new(Pool::new(|| vec![]))),
// CHANGED: Add a flag to have Gitignore rules that apply only to files.
matches: Some(Arc::new(
Pool::with_available_parallelism_capacity(|| vec![]),
)),
// CHANGED: Add a flag to have Gitignore rules that apply only to
// files.
only_on_files: self.only_on_files,
})
}
@ -411,11 +441,8 @@ impl GitignoreBuilder {
// Match Git's handling of .gitignore files that begin with the Unicode BOM
const UTF8_BOM: &str = "\u{feff}";
let line = if i == 0 {
line.trim_start_matches(UTF8_BOM)
} else {
&line
};
let line =
if i == 0 { line.trim_start_matches(UTF8_BOM) } else { &line };
if let Err(err) = self.add_line(Some(path.to_path_buf()), &line) {
errs.push(err.tagged(path, lineno));
@ -537,7 +564,10 @@ impl GitignoreBuilder {
/// affected.
///
/// This is disabled by default.
pub fn case_insensitive(&mut self, yes: bool) -> Result<&mut GitignoreBuilder, Error> {
pub fn case_insensitive(
&mut self,
yes: bool,
) -> Result<&mut GitignoreBuilder, Error> {
// TODO: This should not return a `Result`. Fix this in the next semver
// release.
self.case_insensitive = yes;
@ -556,7 +586,10 @@ impl GitignoreBuilder {
/// modes since the glob parser becomes more permissive. You might want to
/// enable this when compatibility (e.g., with POSIX glob implementations)
/// is more important than good error messages.
pub fn allow_unclosed_class(&mut self, yes: bool) -> &mut GitignoreBuilder {
pub fn allow_unclosed_class(
&mut self,
yes: bool,
) -> &mut GitignoreBuilder {
self.allow_unclosed_class = yes;
self
}
@ -576,32 +609,56 @@ impl GitignoreBuilder {
///
/// Note that the file path returned may not exist.
pub fn gitconfig_excludes_path() -> Option<PathBuf> {
// git supports $HOME/.gitconfig and $XDG_CONFIG_HOME/git/config. Notably,
// both can be active at the same time, where $HOME/.gitconfig takes
// precedent. So if $HOME/.gitconfig defines a `core.excludesFile`, then
// we're done.
match gitconfig_home_contents().and_then(|x| parse_excludes_file(&x)) {
Some(path) => return Some(path),
None => {}
// When GIT_CONFIG_GLOBAL is set, it replaces both $HOME/.gitconfig and
// $XDG_CONFIG_HOME/git/config (per git 2.32+). Otherwise, git supports
// $HOME/.gitconfig and $XDG_CONFIG_HOME/git/config simultaneously, where
// $HOME/.gitconfig takes precedent.
gitconfig_global_env_contents()
.and_then(|x| parse_excludes_file(&x))
.or_else(|| {
gitconfig_home_contents().and_then(|x| parse_excludes_file(&x))
})
.or_else(|| {
gitconfig_xdg_contents().and_then(|x| parse_excludes_file(&x))
})
// System-level config has the lowest priority for core.excludesFile.
// GIT_CONFIG_SYSTEM overrides the default /etc/gitconfig path.
.or_else(|| {
gitconfig_system_contents().and_then(|x| parse_excludes_file(&x))
})
.or_else(excludes_file_default)
}
/// Returns the file contents of git's global config file from the path
/// specified by the `GIT_CONFIG_GLOBAL` environment variable.
fn gitconfig_global_env_contents() -> Option<Vec<u8>> {
let path = std::env::var_os("GIT_CONFIG_GLOBAL").map(PathBuf::from)?;
if path.as_os_str().is_empty() {
return None;
}
match gitconfig_xdg_contents().and_then(|x| parse_excludes_file(&x)) {
Some(path) => return Some(path),
None => {}
}
excludes_file_default()
let mut file = BufReader::new(File::open(path).ok()?);
let mut contents = vec![];
file.read_to_end(&mut contents).ok().map(|_| contents)
}
/// Returns the file contents of git's system-level config file.
///
/// Checks `GIT_CONFIG_SYSTEM` first, then falls back to `/etc/gitconfig`.
fn gitconfig_system_contents() -> Option<Vec<u8>> {
let path = std::env::var_os("GIT_CONFIG_SYSTEM")
.map(PathBuf::from)
.filter(|x| !x.as_os_str().is_empty())
.unwrap_or_else(|| PathBuf::from("/etc/gitconfig"));
let mut file = BufReader::new(File::open(path).ok()?);
let mut contents = vec![];
file.read_to_end(&mut contents).ok().map(|_| contents)
}
/// Returns the file contents of git's global config file, if one exists, in
/// the user's home directory.
fn gitconfig_home_contents() -> Option<Vec<u8>> {
let home = match home_dir() {
None => return None,
Some(home) => home,
};
let mut file = match File::open(home.join(".gitconfig")) {
Err(_) => return None,
Ok(file) => BufReader::new(file),
};
let home = home_dir()?;
let mut file = BufReader::new(File::open(home.join(".gitconfig")).ok()?);
let mut contents = vec![];
file.read_to_end(&mut contents).ok().map(|_| contents)
}
@ -610,19 +667,11 @@ fn gitconfig_home_contents() -> Option<Vec<u8>> {
/// the user's XDG_CONFIG_HOME directory.
fn gitconfig_xdg_contents() -> Option<Vec<u8>> {
let path = std::env::var_os("XDG_CONFIG_HOME")
.and_then(|x| {
if x.is_empty() {
None
} else {
Some(PathBuf::from(x))
}
})
.map(PathBuf::from)
.filter(|x| !x.as_os_str().is_empty())
.or_else(|| home_dir().map(|p| p.join(".config")))
.map(|x| x.join("git/config"));
let mut file = match path.and_then(|p| File::open(p).ok()) {
None => return None,
Some(file) => BufReader::new(file),
};
.map(|x| x.join("git/config"))?;
let mut file = BufReader::new(File::open(path).ok()?);
let mut contents = vec![];
file.read_to_end(&mut contents).ok().map(|_| contents)
}
@ -632,13 +681,8 @@ fn gitconfig_xdg_contents() -> Option<Vec<u8>> {
/// Specifically, this respects XDG_CONFIG_HOME.
fn excludes_file_default() -> Option<PathBuf> {
std::env::var_os("XDG_CONFIG_HOME")
.and_then(|x| {
if x.is_empty() {
None
} else {
Some(PathBuf::from(x))
}
})
.map(PathBuf::from)
.filter(|x| !x.as_os_str().is_empty())
.or_else(|| home_dir().map(|p| p.join(".config")))
.map(|x| x.join("git/ignore"))
}
@ -667,9 +711,7 @@ fn parse_excludes_file(data: &[u8]) -> Option<PathBuf> {
re.captures(data, &mut caps);
let span = caps.get_group(1)?;
let candidate = &data[span];
std::str::from_utf8(candidate)
.ok()
.map(|s| PathBuf::from(expand_tilde(s)))
std::str::from_utf8(candidate).ok().map(|s| PathBuf::from(expand_tilde(s)))
}
/// Expands ~ in file paths to the value of $HOME.
@ -831,7 +873,10 @@ mod tests {
fn parse_excludes_file4() {
let data = bytes("[core]\nexcludesFile = \"~/foo/bar\"");
let got = super::parse_excludes_file(&data);
assert_eq!(path_string(got.unwrap()), super::expand_tilde("~/foo/bar"));
assert_eq!(
path_string(got.unwrap()),
super::expand_tilde("~/foo/bar")
);
}
#[test]

File diff suppressed because it is too large Load diff

View file

@ -48,13 +48,16 @@ See the documentation for `WalkBuilder` for many other options.
use std::path::{Path, PathBuf};
pub use crate::incremental::{IncrementalIgnore, IncrementalMatch};
pub use crate::walk::{
DirEntry, ParallelVisitor, ParallelVisitorBuilder, Walk, WalkBuilder, WalkParallel, WalkState,
DirEntry, ParallelVisitor, ParallelVisitorBuilder, Walk, WalkBuilder,
WalkParallel, WalkState,
};
mod default_types;
mod dir;
pub mod gitignore;
mod incremental;
pub mod overrides;
mod pathutil;
pub mod types;
@ -120,34 +123,31 @@ impl Clone for Error {
fn clone(&self) -> Error {
match *self {
Error::Partial(ref errs) => Error::Partial(errs.clone()),
Error::WithLineNumber { line, ref err } => Error::WithLineNumber {
line,
err: err.clone(),
},
Error::WithPath { ref path, ref err } => Error::WithPath {
path: path.clone(),
err: err.clone(),
},
Error::WithDepth { depth, ref err } => Error::WithDepth {
depth,
err: err.clone(),
},
Error::Loop {
ref ancestor,
ref child,
} => Error::Loop {
Error::WithLineNumber { line, ref err } => {
Error::WithLineNumber { line, err: err.clone() }
}
Error::WithPath { ref path, ref err } => {
Error::WithPath { path: path.clone(), err: err.clone() }
}
Error::WithDepth { depth, ref err } => {
Error::WithDepth { depth, err: err.clone() }
}
Error::Loop { ref ancestor, ref child } => Error::Loop {
ancestor: ancestor.clone(),
child: child.clone(),
},
Error::Io(ref err) => match err.raw_os_error() {
Some(e) => Error::Io(std::io::Error::from_raw_os_error(e)),
None => Error::Io(std::io::Error::new(err.kind(), err.to_string())),
None => {
Error::Io(std::io::Error::new(err.kind(), err.to_string()))
}
},
Error::Glob { ref glob, ref err } => Error::Glob {
glob: glob.clone(),
err: err.clone(),
},
Error::UnrecognizedFileType(ref err) => Error::UnrecognizedFileType(err.clone()),
Error::Glob { ref glob, ref err } => {
Error::Glob { glob: glob.clone(), err: err.clone() }
}
Error::UnrecognizedFileType(ref err) => {
Error::UnrecognizedFileType(err.clone())
}
Error::InvalidDefinition => Error::InvalidDefinition,
}
}
@ -269,19 +269,14 @@ impl Error {
/// Turn an error into a tagged error with the given depth.
fn with_depth(self, depth: usize) -> Error {
Error::WithDepth {
depth,
err: Box::new(self),
}
Error::WithDepth { depth, err: Box::new(self) }
}
/// Turn an error into a tagged error with the given file path and line
/// number. If path is empty, then it is omitted from the error.
fn tagged<P: AsRef<Path>>(self, path: P, lineno: u64) -> Error {
let errline = Error::WithLineNumber {
line: lineno,
err: Box::new(self),
};
let errline =
Error::WithLineNumber { line: lineno, err: Box::new(self) };
if path.as_ref().as_os_str().is_empty() {
return errline;
}
@ -301,12 +296,12 @@ impl Error {
};
}
let path = err.path().map(|p| p.to_path_buf());
let mut ig_err = Error::Io(std::io::Error::from(err));
let mut ig_err = Error::WithDepth {
depth,
err: Box::new(Error::Io(std::io::Error::from(err))),
};
if let Some(path) = path {
ig_err = Error::WithPath {
path,
err: Box::new(ig_err),
};
ig_err = Error::WithPath { path, err: Box::new(ig_err) };
}
ig_err
}
@ -333,7 +328,8 @@ impl std::fmt::Display for Error {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match *self {
Error::Partial(ref errs) => {
let msgs: Vec<String> = errs.iter().map(|err| err.to_string()).collect();
let msgs: Vec<String> =
errs.iter().map(|err| err.to_string()).collect();
write!(f, "{}", msgs.join("\n"))
}
Error::WithLineNumber { line, ref err } => {
@ -343,10 +339,7 @@ impl std::fmt::Display for Error {
write!(f, "{}: {}", path.display(), err)
}
Error::WithDepth { ref err, .. } => err.fmt(f),
Error::Loop {
ref ancestor,
ref child,
} => write!(
Error::Loop { ref ancestor, ref child } => write!(
f,
"File system loop found: \
{} points to an ancestor {}",
@ -354,14 +347,8 @@ impl std::fmt::Display for Error {
ancestor.display()
),
Error::Io(ref err) => err.fmt(f),
Error::Glob {
glob: None,
ref err,
} => write!(f, "{}", err),
Error::Glob {
glob: Some(ref glob),
ref err,
} => {
Error::Glob { glob: None, ref err } => write!(f, "{}", err),
Error::Glob { glob: Some(ref glob), ref err } => {
write!(f, "error parsing glob '{}': {}", glob, err)
}
Error::UnrecognizedFileType(ref ty) => {
@ -507,7 +494,8 @@ mod tests {
};
/// A convenient result type alias.
pub(crate) type Result<T> = std::result::Result<T, Box<dyn std::error::Error + Send + Sync>>;
pub(crate) type Result<T> =
std::result::Result<T, Box<dyn std::error::Error + Send + Sync>>;
macro_rules! err {
($($tt:tt)*) => {
@ -545,8 +533,9 @@ mod tests {
if path.is_dir() {
continue;
}
fs::create_dir_all(&path)
.map_err(|e| err!("failed to create {}: {}", path.display(), e))?;
fs::create_dir_all(&path).map_err(|e| {
err!("failed to create {}: {}", path.display(), e)
})?;
return Ok(TempDir(path));
}
Err(err!("failed to create temp dir after {} tries", TRIES))

View file

@ -94,7 +94,11 @@ impl Override {
/// given) is stripped. If there is no common suffix/prefix overlap, then
/// `path` is assumed to reside in the same directory as the root path for
/// this set of overrides.
pub fn matched<'a, P: AsRef<Path>>(&'a self, path: P, is_dir: bool) -> Match<Glob<'a>> {
pub fn matched<'a, P: AsRef<Path>>(
&'a self,
path: P,
is_dir: bool,
) -> Match<Glob<'a>> {
if self.is_empty() {
return Match::None;
}
@ -146,7 +150,10 @@ impl OverrideBuilder {
/// affected.
///
/// This is disabled by default.
pub fn case_insensitive(&mut self, yes: bool) -> Result<&mut OverrideBuilder, Error> {
pub fn case_insensitive(
&mut self,
yes: bool,
) -> Result<&mut OverrideBuilder, Error> {
// TODO: This should not return a `Result`. Fix this in the next semver
// release.
self.builder.case_insensitive(yes)?;
@ -276,11 +283,8 @@ mod tests {
#[test]
fn default_case_sensitive() {
let ov = OverrideBuilder::new(ROOT)
.add("*.html")
.unwrap()
.build()
.unwrap();
let ov =
OverrideBuilder::new(ROOT).add("*.html").unwrap().build().unwrap();
assert!(ov.matched("foo.html", false).is_whitelist());
assert!(ov.matched("foo.HTML", false).is_ignore());
assert!(ov.matched("foo.htm", false).is_ignore());

View file

@ -2,55 +2,89 @@ use std::{ffi::OsStr, path::Path};
use crate::walk::DirEntry;
/// Returns true if and only if this entry is considered to be hidden.
/// Returns true if and only if this path is considered to be hidden.
///
/// This only returns true if the base name of the path starts with a `.`.
/// # Platform behavior
///
/// On Unix, this implements a more optimized check.
#[cfg(unix)]
pub(crate) fn is_hidden(dent: &DirEntry) -> bool {
use std::os::unix::ffi::OsStrExt;
if let Some(name) = file_name(dent.path()) {
name.as_bytes().get(0) == Some(&b'.')
} else {
false
}
}
/// Returns true if and only if this entry is considered to be hidden.
/// ## Windows
///
/// On Windows, this returns true if one of the following is true:
/// This returns true if one of the following is true:
///
/// * The base name of the path starts with a `.`.
/// * The file attributes have the `HIDDEN` property set.
#[cfg(windows)]
pub(crate) fn is_hidden(dent: &DirEntry) -> bool {
use std::os::windows::fs::MetadataExt;
use winapi_util::file;
// This looks like we're doing an extra stat call, but on Windows, the
// directory traverser reuses the metadata retrieved from each directory
// entry and stores it on the DirEntry itself. So this is "free."
if let Ok(md) = dent.metadata() {
if file::is_hidden(md.file_attributes() as u64) {
return true;
}
}
if let Some(name) = file_name(dent.path()) {
name.to_str().map(|s| s.starts_with(".")).unwrap_or(false)
} else {
false
}
}
/// Returns true if and only if this entry is considered to be hidden.
///
/// ## All other platforms
///
/// This only returns true if the base name of the path starts with a `.`.
#[cfg(not(any(unix, windows)))]
pub(crate) fn is_hidden(dent: &DirEntry) -> bool {
if let Some(name) = file_name(dent.path()) {
name.to_str().map(|s| s.starts_with(".")).unwrap_or(false)
pub(crate) fn is_hidden_path(dent: &Path) -> bool {
#[cfg(not(windows))]
fn imp(path: &Path) -> bool {
is_hidden_path_only(path)
}
#[cfg(windows)]
fn imp(path: &Path) -> bool {
use std::os::windows::fs::MetadataExt;
use winapi_util::file;
if let Ok(md) = path.metadata() {
if file::is_hidden(md.file_attributes() as u64) {
return true;
}
}
is_hidden_path_only(path)
}
imp(dent)
}
/// Returns true if and only if this directory entry is considered to be
/// hidden.
///
/// # Platform behavior
///
/// ## Windows
///
/// This returns true if one of the following is true:
///
/// * The base name of the path starts with a `.`.
/// * The file attributes have the `HIDDEN` property set.
///
/// ## All other platforms
///
/// This only returns true if the base name of the path starts with a `.`.
pub(crate) fn is_hidden_entry(dent: &DirEntry) -> bool {
#[cfg(not(windows))]
fn imp(dent: &DirEntry) -> bool {
is_hidden_path_only(dent.path())
}
#[cfg(windows)]
fn imp(dent: &DirEntry) -> bool {
use std::os::windows::fs::MetadataExt;
use winapi_util::file;
// This looks like we're doing an extra stat call, but on Windows, the
// directory traverser reuses the metadata retrieved from each directory
// entry and stores it on the DirEntry itself. So this is "free."
if let Ok(md) = dent.metadata() {
if file::is_hidden(md.file_attributes() as u64) {
return true;
}
}
is_hidden_path_only(dent.path())
}
imp(dent)
}
/// Returns true if and only if this path is considered to be hidden from only
/// the path itself.
///
/// This has the same behavior on all platforms.
fn is_hidden_path_only(path: &Path) -> bool {
if let Some(name) = file_name(path) {
name.as_encoded_bytes().starts_with(b".")
} else {
false
}
@ -59,83 +93,79 @@ pub(crate) fn is_hidden(dent: &DirEntry) -> bool {
/// Strip `prefix` from the `path` and return the remainder.
///
/// If `path` doesn't have a prefix `prefix`, then return `None`.
#[cfg(unix)]
pub(crate) fn strip_prefix<'a, P: AsRef<Path> + ?Sized>(
prefix: &'a P,
path: &'a Path,
) -> Option<&'a Path> {
use std::os::unix::ffi::OsStrExt;
#[cfg(unix)]
fn imp<'a>(prefix: &'a Path, path: &'a Path) -> Option<&'a Path> {
use std::os::unix::ffi::OsStrExt;
let prefix = prefix.as_ref().as_os_str().as_bytes();
let path = path.as_os_str().as_bytes();
if prefix.len() > path.len() || prefix != &path[0..prefix.len()] {
None
} else {
Some(&Path::new(OsStr::from_bytes(&path[prefix.len()..])))
let prefix = prefix.as_os_str().as_bytes();
let path = path.as_os_str().as_bytes();
if prefix.len() > path.len() || prefix != &path[0..prefix.len()] {
None
} else {
Some(&Path::new(OsStr::from_bytes(&path[prefix.len()..])))
}
}
}
/// Strip `prefix` from the `path` and return the remainder.
///
/// If `path` doesn't have a prefix `prefix`, then return `None`.
#[cfg(not(unix))]
pub(crate) fn strip_prefix<'a, P: AsRef<Path> + ?Sized>(
prefix: &'a P,
path: &'a Path,
) -> Option<&'a Path> {
path.strip_prefix(prefix).ok()
#[cfg(not(unix))]
fn imp<'a>(prefix: &'a Path, path: &'a Path) -> Option<&'a Path> {
path.strip_prefix(prefix).ok()
}
imp(prefix.as_ref(), path)
}
/// Returns true if this file path is just a file name. i.e., Its parent is
/// the empty string.
#[cfg(unix)]
pub(crate) fn is_file_name<P: AsRef<Path>>(path: P) -> bool {
use std::os::unix::ffi::OsStrExt;
use memchr::memchr;
let path = path.as_ref().as_os_str().as_bytes();
memchr(b'/', path).is_none()
}
/// Returns true if this file path is just a file name. i.e., Its parent is
/// the empty string.
#[cfg(not(unix))]
pub(crate) fn is_file_name<P: AsRef<Path>>(path: P) -> bool {
path.as_ref()
.parent()
.map(|p| p.as_os_str().is_empty())
.unwrap_or(false)
}
/// The final component of the path, if it is a normal file.
///
/// If the path terminates in ., .., or consists solely of a root of prefix,
/// file_name will return None.
#[cfg(unix)]
pub(crate) fn file_name<'a, P: AsRef<Path> + ?Sized>(path: &'a P) -> Option<&'a OsStr> {
use memchr::memrchr;
use std::os::unix::ffi::OsStrExt;
let path = path.as_ref().as_os_str().as_bytes();
if path.is_empty() {
return None;
} else if path.len() == 1 && path[0] == b'.' {
return None;
} else if path.last() == Some(&b'.') {
return None;
} else if path.len() >= 2 && &path[path.len() - 2..] == &b".."[..] {
return None;
#[cfg(unix)]
{
memchr::memchr(b'/', path.as_ref().as_os_str().as_encoded_bytes())
.is_none()
}
#[cfg(not(unix))]
{
path.as_ref()
.parent()
.map(|p| p.as_os_str().is_empty())
.unwrap_or(false)
}
let last_slash = memrchr(b'/', path).map(|i| i + 1).unwrap_or(0);
Some(OsStr::from_bytes(&path[last_slash..]))
}
/// The final component of the path, if it is a normal file.
///
/// If the path terminates in ., .., or consists solely of a root of prefix,
/// file_name will return None.
#[cfg(not(unix))]
pub(crate) fn file_name<'a, P: AsRef<Path> + ?Sized>(path: &'a P) -> Option<&'a OsStr> {
path.as_ref().file_name()
/// If the path terminates in `.`, `..`, or consists solely of a root of
/// prefix, this will return `None`.
pub(crate) fn file_name<'a, P: AsRef<Path> + ?Sized>(
path: &'a P,
) -> Option<&'a OsStr> {
#[cfg(unix)]
fn imp(path: &Path) -> Option<&OsStr> {
use std::os::unix::ffi::OsStrExt;
use memchr::memrchr;
let path = path.as_os_str().as_bytes();
if path.is_empty() {
return None;
} else if path.len() == 1 && path[0] == b'.' {
return None;
} else if path.last() == Some(&b'.') {
return None;
} else if path.len() >= 2 && &path[path.len() - 2..] == &b".."[..] {
return None;
}
let last_slash = memrchr(b'/', path).map(|i| i + 1).unwrap_or(0);
Some(OsStr::from_bytes(&path[last_slash..]))
}
#[cfg(not(unix))]
fn imp(path: &Path) -> Option<&OsStr> {
path.file_name()
}
imp(path.as_ref())
}

View file

@ -204,8 +204,12 @@ impl<T> Selection<T> {
fn map<U, F: FnOnce(T) -> U>(self, f: F) -> Selection<U> {
match self {
Selection::Select(name, inner) => Selection::Select(name, f(inner)),
Selection::Negate(name, inner) => Selection::Negate(name, f(inner)),
Selection::Select(name, inner) => {
Selection::Select(name, f(inner))
}
Selection::Negate(name, inner) => {
Selection::Negate(name, f(inner))
}
}
}
@ -227,7 +231,9 @@ impl Types {
has_selected: false,
glob_to_selection: vec![],
set: GlobSetBuilder::new().build().unwrap(),
matches: Arc::new(Pool::new(|| vec![])),
matches: Arc::new(Pool::with_available_parallelism_capacity(
|| vec![],
)),
}
}
@ -254,7 +260,11 @@ impl Types {
/// The path is considered ignored if it matches a negated file type.
/// If at least one file type is selected and `path` doesn't match, then
/// the path is also considered ignored.
pub fn matched<'a, P: AsRef<Path>>(&'a self, path: P, is_dir: bool) -> Match<Glob<'a>> {
pub fn matched<'a, P: AsRef<Path>>(
&'a self,
path: P,
is_dir: bool,
) -> Match<Glob<'a>> {
// File types don't apply to directories, and we can't do anything
// if our glob set is empty.
if is_dir || self.set.is_empty() {
@ -306,10 +316,7 @@ impl TypesBuilder {
/// of default type definitions can be added with `add_defaults`, and
/// additional type definitions can be added with `select` and `negate`.
pub fn new() -> TypesBuilder {
TypesBuilder {
types: HashMap::new(),
selections: vec![],
}
TypesBuilder { types: HashMap::new(), selections: vec![] }
}
/// Build the current set of file type definitions *and* selections into
@ -343,17 +350,18 @@ impl TypesBuilder {
}
selections.push(selection.clone().map(move |_| def));
}
let set = build_set.build().map_err(|err| Error::Glob {
glob: None,
err: err.to_string(),
})?;
let set = build_set
.build()
.map_err(|err| Error::Glob { glob: None, err: err.to_string() })?;
Ok(Types {
defs,
selections,
has_selected,
glob_to_selection,
set,
matches: Arc::new(Pool::new(|| vec![])),
matches: Arc::new(Pool::with_available_parallelism_capacity(
|| vec![],
)),
})
}
@ -377,12 +385,10 @@ impl TypesBuilder {
pub fn select(&mut self, name: &str) -> &mut TypesBuilder {
if name == "all" {
for name in self.types.keys() {
self.selections
.push(Selection::Select(name.to_string(), ()));
self.selections.push(Selection::Select(name.to_string(), ()));
}
} else {
self.selections
.push(Selection::Select(name.to_string(), ()));
self.selections.push(Selection::Select(name.to_string(), ()));
}
self
}
@ -393,12 +399,10 @@ impl TypesBuilder {
pub fn negate(&mut self, name: &str) -> &mut TypesBuilder {
if name == "all" {
for name in self.types.keys() {
self.selections
.push(Selection::Negate(name.to_string(), ()));
self.selections.push(Selection::Negate(name.to_string(), ()));
}
} else {
self.selections
.push(Selection::Negate(name.to_string(), ()));
self.selections.push(Selection::Negate(name.to_string(), ()));
}
self
}
@ -453,7 +457,10 @@ impl TypesBuilder {
3 => {
let name = parts[0];
let types_string = parts[2];
if name.is_empty() || parts[1] != "include" || types_string.is_empty() {
if name.is_empty()
|| parts[1] != "include"
|| types_string.is_empty()
{
return Err(Error::InvalidDefinition);
}
let types = types_string.split(',');
@ -463,7 +470,8 @@ impl TypesBuilder {
return Err(Error::InvalidDefinition);
}
for type_name in types {
let globs = self.types.get(type_name).unwrap().globs.clone();
let globs =
self.types.get(type_name).unwrap().globs.clone();
for glob in globs {
self.add(name, &glob)?;
}
@ -549,30 +557,9 @@ mod tests {
matched!(not, matchnot1, types(), vec!["rust"], vec![], "index.html");
matched!(not, matchnot2, types(), vec![], vec!["rust"], "main.rs");
matched!(
not,
matchnot3,
types(),
vec!["foo"],
vec!["rust"],
"main.rs"
);
matched!(
not,
matchnot4,
types(),
vec!["rust"],
vec!["foo"],
"main.rs"
);
matched!(
not,
matchnot5,
types(),
vec!["rust"],
vec!["foo"],
"main.foo"
);
matched!(not, matchnot3, types(), vec!["foo"], vec!["rust"], "main.rs");
matched!(not, matchnot4, types(), vec!["rust"], vec!["foo"], "main.rs");
matched!(not, matchnot5, types(), vec!["rust"], vec!["foo"], "main.foo");
matched!(not, matchnot6, types(), vec!["combo"], vec![], "leftpad.js");
matched!(not, matchnot7, types(), vec!["py"], vec![], "index.html");
matched!(not, matchnot8, types(), vec!["python"], vec![], "doc.md");

View file

@ -17,7 +17,9 @@ use {
use crate::{
Error, PartialErrorBuilder,
dir::{Ignore, IgnoreBuilder},
// CHANGED: Also import `Gitignore` for `WalkBuilder::add_gitignore`.
gitignore::{Gitignore, GitignoreBuilder},
incremental::{IncrementalIgnore, IncrementalIgnoreOptions},
overrides::Override,
types::Types,
};
@ -104,24 +106,15 @@ impl DirEntry {
}
fn new_stdin() -> DirEntry {
DirEntry {
dent: DirEntryInner::Stdin,
err: None,
}
DirEntry { dent: DirEntryInner::Stdin, err: None }
}
fn new_walkdir(dent: walkdir::DirEntry, err: Option<Error>) -> DirEntry {
DirEntry {
dent: DirEntryInner::Walkdir(dent),
err,
}
DirEntry { dent: DirEntryInner::Walkdir(dent), err }
}
fn new_raw(dent: DirEntryRaw, err: Option<Error>) -> DirEntry {
DirEntry {
dent: DirEntryInner::Raw(dent),
err,
}
DirEntry { dent: DirEntryInner::Raw(dent), err }
}
}
@ -187,9 +180,11 @@ impl DirEntryInner {
));
Err(err.with_path("<stdin>"))
}
Walkdir(ref x) => x
.metadata()
.map_err(|err| Error::Io(io::Error::from(err)).with_path(x.path())),
Walkdir(ref x) => x.metadata().map_err(|err| {
Error::Io(io::Error::from(err))
.with_depth(x.depth())
.with_path(x.path())
}),
Raw(ref x) => x.metadata(),
}
}
@ -308,7 +303,9 @@ impl DirEntryRaw {
} else {
fs::symlink_metadata(&self.path)
}
.map_err(|err| Error::Io(io::Error::from(err)).with_path(&self.path))
.map_err(|err| {
Error::Io(err).with_depth(self.depth).with_path(&self.path)
})
}
fn file_type(&self) -> FileType {
@ -316,9 +313,7 @@ impl DirEntryRaw {
}
fn file_name(&self) -> &OsStr {
self.path
.file_name()
.unwrap_or_else(|| self.path.as_os_str())
self.path.file_name().unwrap_or_else(|| self.path.as_os_str())
}
fn depth(&self) -> usize {
@ -330,13 +325,13 @@ impl DirEntryRaw {
self.ino
}
fn from_entry(depth: usize, ent: &fs::DirEntry) -> Result<DirEntryRaw, Error> {
fn from_entry(
depth: usize,
ent: &fs::DirEntry,
) -> Result<DirEntryRaw, Error> {
let ty = ent.file_type().map_err(|err| {
let err = Error::Io(io::Error::from(err)).with_path(ent.path());
Error::WithDepth {
depth,
err: Box::new(err),
}
let err = Error::Io(err).with_depth(depth).with_path(ent.path());
Error::WithDepth { depth, err: Box::new(err) }
})?;
DirEntryRaw::from_entry_os(depth, ent, ty)
}
@ -348,11 +343,8 @@ impl DirEntryRaw {
ty: fs::FileType,
) -> Result<DirEntryRaw, Error> {
let md = ent.metadata().map_err(|err| {
let err = Error::Io(io::Error::from(err)).with_path(ent.path());
Error::WithDepth {
depth,
err: Box::new(err),
}
let err = Error::Io(err).with_depth(depth).with_path(ent.path());
Error::WithDepth { depth, err: Box::new(err) }
})?;
Ok(DirEntryRaw {
path: ent.path(),
@ -395,8 +387,13 @@ impl DirEntryRaw {
}
#[cfg(windows)]
fn from_path(depth: usize, pb: PathBuf, link: bool) -> Result<DirEntryRaw, Error> {
let md = fs::metadata(&pb).map_err(|err| Error::Io(err).with_path(&pb))?;
fn from_path(
depth: usize,
pb: PathBuf,
link: bool,
) -> Result<DirEntryRaw, Error> {
let md = fs::metadata(&pb)
.map_err(|err| Error::Io(err).with_depth(depth).with_path(&pb))?;
Ok(DirEntryRaw {
path: pb,
ty: md.file_type(),
@ -407,10 +404,15 @@ impl DirEntryRaw {
}
#[cfg(unix)]
fn from_path(depth: usize, pb: PathBuf, link: bool) -> Result<DirEntryRaw, Error> {
fn from_path(
depth: usize,
pb: PathBuf,
link: bool,
) -> Result<DirEntryRaw, Error> {
use std::os::unix::fs::MetadataExt;
let md = fs::metadata(&pb).map_err(|err| Error::Io(err).with_path(&pb))?;
let md = fs::metadata(&pb)
.map_err(|err| Error::Io(err).with_depth(depth).with_path(&pb))?;
Ok(DirEntryRaw {
path: pb,
ty: md.file_type(),
@ -423,7 +425,11 @@ impl DirEntryRaw {
// Placeholder implementation to allow compiling on non-standard platforms
// (e.g. wasm32).
#[cfg(not(any(windows, unix)))]
fn from_path(depth: usize, pb: PathBuf, link: bool) -> Result<DirEntryRaw, Error> {
fn from_path(
depth: usize,
pb: PathBuf,
link: bool,
) -> Result<DirEntryRaw, Error> {
Err(Error::Io(io::Error::new(
io::ErrorKind::Other,
"unsupported platform",
@ -502,7 +508,8 @@ pub struct WalkBuilder {
///
/// When `None`, the CWD is fetched from `std::env::current_dir()`. If
/// that fails, then global gitignores are ignored (an error is logged).
global_gitignores_relative_to: OnceLock<Result<PathBuf, Arc<std::io::Error>>>,
global_gitignores_relative_to:
OnceLock<Result<PathBuf, Arc<std::io::Error>>>,
}
#[derive(Clone)]
@ -544,8 +551,16 @@ impl WalkBuilder {
/// is better to call `add` on this builder than to create multiple
/// `Walk` values.
pub fn new<P: AsRef<Path>>(path: P) -> WalkBuilder {
WalkBuilder::from_iter([path])
}
/// Create an empty builder to which paths can be added.
///
/// Note that if you call `build` on this instance before calling `add`
/// on it, it will return exactly zero items during iteration.
pub fn empty() -> WalkBuilder {
WalkBuilder {
paths: vec![path.as_ref().to_path_buf()],
paths: vec![],
ig_builder: IgnoreBuilder::new(),
max_depth: None,
min_depth: None,
@ -560,6 +575,21 @@ impl WalkBuilder {
}
}
/// Create a new builder for a recursive directory iterator from the
/// sequence of paths.
///
/// Note that if the iterator is empty, this is the same as
/// `WalkBuilder::empty`.
pub fn from_iter<P: AsRef<Path>, I: IntoIterator<Item = P>>(
paths: I,
) -> WalkBuilder {
let mut builder = WalkBuilder::empty();
for path in paths.into_iter() {
builder.add(path);
}
builder
}
/// Build a new `Walk` iterator.
pub fn build(&self) -> Walk {
let follow_links = self.follow_links;
@ -585,10 +615,14 @@ impl WalkBuilder {
if let Some(ref sorter) = sorter {
match sorter.clone() {
Sorter::ByName(cmp) => {
wd = wd.sort_by(move |a, b| cmp(a.file_name(), b.file_name()));
wd = wd.sort_by(move |a, b| {
cmp(a.file_name(), b.file_name())
});
}
Sorter::ByPath(cmp) => {
wd = wd.sort_by(move |a, b| cmp(a.path(), b.path()));
wd = wd.sort_by(move |a, b| {
cmp(a.path(), b.path())
});
}
}
}
@ -597,31 +631,76 @@ impl WalkBuilder {
})
.collect::<Vec<_>>()
.into_iter();
let ig_root = self
.get_or_set_current_dir()
.map(|cwd| self.ig_builder.build_with_cwd(Some(cwd.to_path_buf())))
.unwrap_or_else(|| self.ig_builder.build());
let ig_root = self.build_ignore();
Walk {
its,
it: None,
ig_root: ig_root.clone(),
ig: ig_root.clone(),
max_depth: self.max_depth,
max_filesize: self.max_filesize,
skip: self.skip.clone(),
filter: self.filter.clone(),
}
}
/// Build matchers for checking paths against ignore files without
/// recursively walking the configured roots.
///
/// The returned matchers use the path-based filtering configuration
/// on this builder, including glob overrides, file type selections,
/// parent ignore files, `.ignore`, `.gitignore`, global Git
/// ignore files, explicitly added ignore files and custom ignore
/// file names. For example, ripgrep configures `.rgignore` via
/// [`WalkBuilder::add_custom_ignore_filename`]. Minimum and maximum depth
/// limits, maximum file size and hidden-file filtering are also applied.
/// Other options that only control traversal or require a directory entry,
/// such as custom entry predicates, are not applied.
///
/// One matcher is returned for each configured path, in the same order as
/// the paths were added to this builder. Each matcher accepts paths
/// relative to its own [`IncrementalIgnore::root`]. The matcher for the
/// special `-` path representing standard input always returns a non-match
/// for all inputs.
///
/// Ignore matchers are loaded lazily and cached by directory.
/// Thus, the first query may read ignore files from the root and
/// its parents, while later queries reuse the compiled matchers.
/// Errors encountered while loading ignore files are returned by
/// [`IncrementalIgnore::matched_with_errors`]. Once an ignore file has
/// been loaded, changes to it are not observed. Build new matchers to
/// reload changed ignore files.
///
/// Matchers built together share the builder's base ignore configuration
/// and compiled parent matchers.
pub fn build_matchers(&self) -> Vec<IncrementalIgnore> {
let ignore = self.build_ignore();
let options = IncrementalIgnoreOptions {
min_depth: self.min_depth,
max_depth: self.max_depth,
max_filesize: self.max_filesize,
hidden: self.ig_builder.is_hidden(),
follow_links: self.follow_links,
};
self.paths
.iter()
.map(move |path| {
IncrementalIgnore::new(
path.clone(),
ignore.clone(),
options.clone(),
)
})
.collect()
}
/// Build a new `WalkParallel` iterator.
///
/// Note that this *doesn't* return something that implements `Iterator`.
/// Instead, the returned value must be run with a closure. e.g.,
/// `builder.build_parallel().run(|| |path| { println!("{path:?}"); WalkState::Continue })`.
pub fn build_parallel(&self) -> WalkParallel {
let ig_root = self
.get_or_set_current_dir()
.map(|cwd| self.ig_builder.build_with_cwd(Some(cwd.to_path_buf())))
.unwrap_or_else(|| self.ig_builder.build());
let ig_root = self.build_ignore();
WalkParallel {
paths: self.paths.clone().into_iter(),
ig_root,
@ -651,7 +730,10 @@ impl WalkBuilder {
/// The default, `None`, imposes no depth restriction.
pub fn max_depth(&mut self, depth: Option<usize>) -> &mut WalkBuilder {
self.max_depth = depth;
if self.min_depth.is_some() && self.max_depth.is_some() && self.max_depth < self.min_depth {
if self.min_depth.is_some()
&& self.max_depth.is_some()
&& self.max_depth < self.min_depth
{
self.max_depth = self.min_depth;
}
self
@ -662,7 +744,10 @@ impl WalkBuilder {
/// The default, `None`, imposes no minimum depth restriction.
pub fn min_depth(&mut self, depth: Option<usize>) -> &mut WalkBuilder {
self.min_depth = depth;
if self.max_depth.is_some() && self.min_depth.is_some() && self.min_depth > self.max_depth {
if self.max_depth.is_some()
&& self.min_depth.is_some()
&& self.min_depth > self.max_depth
{
self.min_depth = self.max_depth;
}
self
@ -705,7 +790,12 @@ impl WalkBuilder {
/// An error will also occur if this walker could not get the current
/// working directory (and `WalkBuilder::current_dir` isn't set).
pub fn add_ignore<P: AsRef<Path>>(&mut self, path: P) -> Option<Error> {
// CHANGED: Dropped this code
// CHANGED: Root the ignore file at `""` instead of the current working
// directory. Explicit ignores are scoped to the directory of the
// ignore file (see `matched_ignore`), and a root of `""` makes the
// rules apply to every walked path regardless of the walk root. This
// also avoids depending on the current working directory entirely.
//
// let path = path.as_ref();
// let Some(cwd) = self.get_or_set_current_dir() else {
// let err = std::io::Error::other(format!(
@ -729,7 +819,11 @@ impl WalkBuilder {
errs.into_error_option()
}
/// CHANGED: Add a Gitignore to the builder.
/// CHANGED: Add a prebuilt Gitignore to the builder.
///
/// Like the ignore file added via `add_ignore`, these rules are matched
/// against the full path of each walked entry, scoped to the `Gitignore`'s
/// root path.
pub fn add_gitignore(&mut self, gi: Gitignore) {
self.ig_builder.add_ignore(gi);
}
@ -982,7 +1076,10 @@ impl WalkBuilder {
///
/// Global gitignore files come from things like a user's git configuration
/// or from gitignore files added via [`WalkBuilder::add_ignore`].
pub fn current_dir(&mut self, cwd: impl Into<PathBuf>) -> &mut WalkBuilder {
pub fn current_dir(
&mut self,
cwd: impl Into<PathBuf>,
) -> &mut WalkBuilder {
let cwd = cwd.into();
self.ig_builder.current_dir(cwd.clone());
if let Err(cwd) = self.global_gitignores_relative_to.set(Ok(cwd)) {
@ -1002,7 +1099,10 @@ impl WalkBuilder {
let result = std::env::current_dir().map_err(Arc::new);
match result {
Ok(ref path) => {
log::trace!("automatically discovered CWD: {}", path.display());
log::trace!(
"automatically discovered CWD: {}",
path.display()
);
}
Err(ref err) => {
log::debug!(
@ -1016,6 +1116,13 @@ impl WalkBuilder {
});
result.as_ref().ok().map(|path| &**path)
}
/// Build the root ignore matcher shared by all consumers of this builder.
fn build_ignore(&self) -> Ignore {
self.get_or_set_current_dir()
.map(|cwd| self.ig_builder.build_with_cwd(Some(cwd.to_path_buf())))
.unwrap_or_else(|| self.ig_builder.build())
}
}
/// Walk is a recursive directory iterator over file paths in one or more
@ -1029,6 +1136,7 @@ pub struct Walk {
it: Option<WalkEventIter>,
ig_root: Ignore,
ig: Ignore,
max_depth: Option<usize>,
max_filesize: Option<u64>,
skip: Option<Arc<Handle>>,
filter: Option<Filter>,
@ -1044,6 +1152,17 @@ impl Walk {
WalkBuilder::new(path).build()
}
/// Create a new recursive directory iterator from the sequence of paths
/// given.
///
/// Note that if the provided iterator is empty, then `Walk` is guaranteed
/// to yield zero entries.
pub fn from_iter<P: AsRef<Path>, I: IntoIterator<Item = P>>(
paths: I,
) -> Walk {
WalkBuilder::from_iter(paths).build()
}
fn skip_entry(&self, ent: &DirEntry) -> Result<bool, Error> {
if ent.depth() == 0 {
return Ok(false);
@ -1128,12 +1247,17 @@ impl Iterator for Walk {
self.it.as_mut().unwrap().it.skip_current_dir();
// Still need to push this on the stack because
// we'll get a WalkEvent::Exit event for this dir.
// We don't care if it errors though.
let (igtmp, _) = self.ig.add_child(ent.path());
// Its ignore files cannot apply to any visited entry.
let (igtmp, _) =
self.ig.add_child_with_entries(ent.path(), &[]);
self.ig = igtmp;
continue;
}
let (igtmp, err) = self.ig.add_child(ent.path());
let (igtmp, err) = if self.max_depth == Some(ent.depth()) {
self.ig.add_child_with_entries(ent.path(), &[])
} else {
self.ig.add_child(ent.path())
};
self.ig = igtmp;
ent.err = err;
return Some(Ok(ent));
@ -1175,11 +1299,7 @@ enum WalkEvent {
impl From<WalkDir> for WalkEventIter {
fn from(it: WalkDir) -> WalkEventIter {
WalkEventIter {
depth: 0,
it: it.into_iter(),
next: None,
}
WalkEventIter { depth: 0, it: it.into_iter(), next: None }
}
}
@ -1252,7 +1372,9 @@ pub trait ParallelVisitorBuilder<'s> {
fn build(&mut self) -> Box<dyn ParallelVisitor + 's>;
}
impl<'a, 's, P: ParallelVisitorBuilder<'s>> ParallelVisitorBuilder<'s> for &'a mut P {
impl<'a, 's, P: ParallelVisitorBuilder<'s>> ParallelVisitorBuilder<'s>
for &'a mut P
{
fn build(&mut self) -> Box<dyn ParallelVisitor + 's> {
(**self).build()
}
@ -1273,14 +1395,17 @@ struct FnBuilder<F> {
builder: F,
}
impl<'s, F: FnMut() -> FnVisitor<'s>> ParallelVisitorBuilder<'s> for FnBuilder<F> {
impl<'s, F: FnMut() -> FnVisitor<'s>> ParallelVisitorBuilder<'s>
for FnBuilder<F>
{
fn build(&mut self) -> Box<dyn ParallelVisitor + 's> {
let visitor = (self.builder)();
Box::new(FnVisitorImp { visitor })
}
}
type FnVisitor<'s> = Box<dyn FnMut(Result<DirEntry, Error>) -> WalkState + Send + 's>;
type FnVisitor<'s> =
Box<dyn FnMut(Result<DirEntry, Error>) -> WalkState + Send + 's>;
struct FnVisitorImp<'s> {
visitor: FnVisitor<'s>,
@ -1370,7 +1495,9 @@ impl WalkParallel {
}
};
match DirEntryRaw::from_path(0, path, false) {
Ok(dent) => (DirEntry::new_raw(dent, None), root_device),
Ok(dent) => {
(DirEntry::new_raw(dent, None), root_device)
}
Err(err) => {
if visitor.visit(Err(err)).is_quit() {
return;
@ -1394,21 +1521,28 @@ impl WalkParallel {
let quit_now = Arc::new(AtomicBool::new(false));
let active_workers = Arc::new(AtomicUsize::new(threads));
let stacks = Stack::new_for_each_thread(threads, stack);
// Collect all of the workers first. In the case that
// `builder.build()` panics, we want that to happen and
// propagate before we actually start to run any of the
// workers.
let workers: Vec<_> = stacks
.into_iter()
.map(|stack| Worker {
visitor: builder.build(),
stack,
quit_now: quit_now.clone(),
active_workers: active_workers.clone(),
max_depth: self.max_depth,
min_depth: self.min_depth,
max_filesize: self.max_filesize,
follow_links: self.follow_links,
skip: self.skip.clone(),
filter: self.filter.clone(),
})
.collect();
std::thread::scope(|s| {
let handles: Vec<_> = stacks
let handles: Vec<_> = workers
.into_iter()
.map(|stack| Worker {
visitor: builder.build(),
stack,
quit_now: quit_now.clone(),
active_workers: active_workers.clone(),
max_depth: self.max_depth,
min_depth: self.min_depth,
max_filesize: self.max_filesize,
follow_links: self.follow_links,
skip: self.skip.clone(),
filter: self.filter.clone(),
})
.map(|worker| s.spawn(|| worker.run()))
.collect();
for handle in handles {
@ -1419,9 +1553,7 @@ impl WalkParallel {
fn threads(&self) -> usize {
if self.threads == 0 {
std::thread::available_parallelism()
.map_or(1, |n| n.get())
.min(12)
std::thread::available_parallelism().map_or(1, |n| n.get()).min(12)
} else {
self.threads
}
@ -1452,6 +1584,12 @@ struct Work {
root_device: Option<u64>,
}
#[derive(Default)]
struct ReadDirResult {
entries: Vec<fs::DirEntry>,
errors: Vec<Error>,
}
impl Work {
/// Returns true if and only if this work item is a directory.
fn is_dir(&self) -> bool {
@ -1478,6 +1616,13 @@ impl Work {
err
}
/// Adds ignore rules for this directory without reading its contents.
fn add_ignore(&mut self) {
let (ig, err) = self.ignore.add_child(self.dent.path());
self.ignore = ig;
self.dent.err = err;
}
/// Reads the directory contents of this work item and adds ignore
/// rules for this directory.
///
@ -1485,7 +1630,7 @@ impl Work {
/// an error is returned. If there was a problem reading the ignore
/// rules for this directory, then the error is attached to this
/// work item's directory entry.
fn read_dir(&mut self) -> Result<fs::ReadDir, Error> {
fn read_dir(&mut self) -> Result<ReadDirResult, Error> {
let readdir = match fs::read_dir(self.dent.path()) {
Ok(readdir) => readdir,
Err(err) => {
@ -1495,10 +1640,24 @@ impl Work {
return Err(err);
}
};
let (ig, err) = self.ignore.add_child(self.dent.path());
// Actually descend into the directory and read its contents
let mut result = ReadDirResult::default();
for entry in readdir {
match entry {
Ok(entry) => result.entries.push(entry),
Err(err) => result.errors.push(
Error::from(err)
.with_path(self.dent.path())
.with_depth(self.dent.depth() + 1),
),
}
}
let (ig, err) = self
.ignore
.add_child_with_entries(self.dent.path(), &result.entries);
self.ignore = ig;
self.dent.err = err;
Ok(readdir)
Ok(result)
}
}
@ -1522,11 +1681,11 @@ impl Stack {
// breadth-first. We do depth-first because a breadth first traversal
// on wide directories with a lot of gitignores is disastrous (for
// example, searching a directory tree containing all of crates.io).
let deques: Vec<Deque<Message>> = std::iter::repeat_with(Deque::new_lifo)
.take(threads)
.collect();
let stealers =
Arc::<[Stealer<Message>]>::from(deques.iter().map(Deque::stealer).collect::<Vec<_>>());
let deques: Vec<Deque<Message>> =
std::iter::repeat_with(Deque::new_lifo).take(threads).collect();
let stealers = Arc::<[Stealer<Message>]>::from(
deques.iter().map(Deque::stealer).collect::<Vec<_>>(),
);
let stacks: Vec<Stack> = deques
.into_iter()
.enumerate()
@ -1668,8 +1827,13 @@ impl<'s> Worker<'s> {
// have sufficient read permissions to list the directory.
// In that case we still want to provide the closure with a valid
// entry before passing the error value.
let readdir = work.read_dir();
let depth = work.dent.depth();
let readdir = if descend && self.max_depth.is_none_or(|m| depth < m) {
Some(work.read_dir())
} else {
work.add_ignore();
None
};
if should_visit {
let state = self.visitor.visit(Ok(work.dent));
if !state.is_continue() {
@ -1680,6 +1844,10 @@ impl<'s> Worker<'s> {
return WalkState::Skip;
}
let readdir = match readdir {
Some(readdir) => readdir,
None => return WalkState::Skip,
};
let readdir = match readdir {
Ok(readdir) => readdir,
Err(err) => {
@ -1687,11 +1855,19 @@ impl<'s> Worker<'s> {
}
};
if self.max_depth.map_or(false, |max| depth >= max) {
return WalkState::Skip;
for result in readdir.entries {
let state = self.generate_work(
&work.ignore,
depth + 1,
work.root_device,
result,
);
if state.is_quit() {
return state;
}
}
for result in readdir {
let state = self.generate_work(&work.ignore, depth + 1, work.root_device, result);
for err in readdir.errors {
let state = self.visitor.visit(Err(err));
if state.is_quit() {
return state;
}
@ -1717,14 +1893,8 @@ impl<'s> Worker<'s> {
ig: &Ignore,
depth: usize,
root_device: Option<u64>,
result: Result<fs::DirEntry, io::Error>,
fs_dent: fs::DirEntry,
) -> WalkState {
let fs_dent = match result {
Ok(fs_dent) => fs_dent,
Err(err) => {
return self.visitor.visit(Err(Error::from(err).with_depth(depth)));
}
};
let mut dent = match DirEntryRaw::from_entry(depth, &fs_dent) {
Ok(dent) => DirEntry::new_raw(dent, None),
Err(err) => {
@ -1760,26 +1930,24 @@ impl<'s> Worker<'s> {
return WalkState::Continue;
}
}
let should_skip_filesize = if self.max_filesize.is_some() && !dent.is_dir() {
skip_filesize(
self.max_filesize.unwrap(),
dent.path(),
&dent.metadata().ok(),
)
} else {
false
};
let should_skip_filtered = if let Some(Filter(predicate)) = &self.filter {
!predicate(&dent)
} else {
false
};
let should_skip_filesize =
if self.max_filesize.is_some() && !dent.is_dir() {
skip_filesize(
self.max_filesize.unwrap(),
dent.path(),
&dent.metadata().ok(),
)
} else {
false
};
let should_skip_filtered =
if let Some(Filter(predicate)) = &self.filter {
!predicate(&dent)
} else {
false
};
if !should_skip_filesize && !should_skip_filtered {
self.send(Work {
dent,
ignore: ig.clone(),
root_device,
});
self.send(Work { dent, ignore: ig.clone(), root_device });
}
WalkState::Continue
}
@ -1820,6 +1988,9 @@ impl<'s> Worker<'s> {
}
// Wait for next `Work` or `Quit` message.
loop {
if self.is_quit_now() {
return None;
}
if let Some(v) = self.recv() {
self.activate_worker();
value = Some(v);
@ -1873,24 +2044,25 @@ impl<'s> Worker<'s> {
}
}
impl<'s> Drop for Worker<'s> {
fn drop(&mut self) {
if std::thread::panicking() {
self.quit_now();
}
}
}
fn check_symlink_loop(
ig_parent: &Ignore,
child_path: &Path,
child_depth: usize,
) -> Result<(), Error> {
let hchild = Handle::from_path(child_path).map_err(|err| {
Error::from(err)
.with_path(child_path)
.with_depth(child_depth)
Error::from(err).with_path(child_path).with_depth(child_depth)
})?;
for ig in ig_parent
.parents()
.take_while(|ig| !ig.is_absolute_parent())
{
for ig in ig_parent.parents().take_while(|ig| !ig.is_absolute_parent()) {
let h = Handle::from_path(ig.path()).map_err(|err| {
Error::from(err)
.with_path(child_path)
.with_depth(child_depth)
Error::from(err).with_path(child_path).with_depth(child_depth)
})?;
if hchild == h {
return Err(Error::Loop {
@ -1905,7 +2077,11 @@ fn check_symlink_loop(
// Before calling this function, make sure that you ensure that is really
// necessary as the arguments imply a file stat.
fn skip_filesize(max_filesize: u64, path: &Path, ent: &Option<Metadata>) -> bool {
fn skip_filesize(
max_filesize: u64,
path: &Path,
ent: &Option<Metadata>,
) -> bool {
let filesize = match *ent {
Some(ref md) => Some(md.len()),
None => None,
@ -1977,9 +2153,9 @@ fn path_equals(dent: &DirEntry, handle: &Handle) -> Result<bool, Error> {
if dent.is_stdin() || never_equal(dent, handle) {
return Ok(false);
}
Handle::from_path(dent.path())
.map(|h| &h == handle)
.map_err(|err| Error::Io(err).with_path(dent.path()))
Handle::from_path(dent.path()).map(|h| &h == handle).map_err(|err| {
Error::Io(err).with_depth(dent.depth()).with_path(dent.path())
})
}
/// Returns true if the given walkdir entry corresponds to a directory.
@ -1997,16 +2173,14 @@ fn walkdir_is_dir(dent: &walkdir::DirEntry) -> bool {
if !dent.file_type().is_symlink() || dent.depth() > 0 {
return false;
}
dent.path()
.metadata()
.ok()
.map_or(false, |md| md.file_type().is_dir())
dent.path().metadata().ok().map_or(false, |md| md.file_type().is_dir())
}
/// Returns true if and only if the given path is on the same device as the
/// given root device.
fn is_same_file_system(root_device: u64, path: &Path) -> Result<bool, Error> {
let dent_device = device_num(path).map_err(|err| Error::Io(err).with_path(path))?;
let dent_device =
device_num(path).map_err(|err| Error::Io(err).with_path(path))?;
Ok(root_device == dent_device)
}
@ -2065,11 +2239,7 @@ mod tests {
}
fn normal_path(unix: &str) -> String {
if cfg!(windows) {
unix.replace("\\", "/")
} else {
unix.to_string()
}
if cfg!(windows) { unix.replace("\\", "/") } else { unix.to_string() }
}
fn walk_collect(prefix: &Path, builder: &WalkBuilder) -> Vec<String> {
@ -2089,7 +2259,10 @@ mod tests {
paths
}
fn walk_collect_parallel(prefix: &Path, builder: &WalkBuilder) -> Vec<String> {
fn walk_collect_parallel(
prefix: &Path,
builder: &WalkBuilder,
) -> Vec<String> {
let mut paths = vec![];
for dent in walk_collect_entries_parallel(builder) {
let path = dent.path().strip_prefix(prefix).unwrap();
@ -2278,6 +2451,27 @@ mod tests {
);
}
#[test]
fn max_depth_does_not_load_unreachable_ignore_files() {
let td = tmpdir();
let leaf = td.path().join("leaf");
mkdirp(&leaf);
wfile(leaf.join(".ignore"), "{invalid\n");
let mut builder = WalkBuilder::new(td.path());
builder.max_depth(Some(1));
let entry = builder
.build()
.find_map(|result| {
let entry = result.unwrap();
(entry.path() == leaf).then_some(entry)
})
.unwrap();
assert!(entry.error().is_none());
assert_paths(td.path(), &builder, &["leaf"]);
}
#[test]
fn min_depth() {
let td = tmpdir();
@ -2388,7 +2582,9 @@ mod tests {
assert_eq!(1, dents.len());
assert!(!dents[0].path_is_symlink());
let dents = walk_collect_entries_parallel(&WalkBuilder::new(td.path().join("foo")));
let dents = walk_collect_entries_parallel(&WalkBuilder::new(
td.path().join("foo"),
));
assert_eq!(1, dents.len());
assert!(!dents[0].path_is_symlink());
}
@ -2474,8 +2670,88 @@ mod tests {
assert_paths(
td.path(),
&WalkBuilder::new(td.path()).filter_entry(|entry| entry.file_name() != OsStr::new("a")),
&WalkBuilder::new(td.path())
.filter_entry(|entry| entry.file_name() != OsStr::new("a")),
&["x", "x/y", "x/y/foo"],
);
}
#[test]
fn empty() {
let td = tmpdir();
assert_paths(td.path(), &WalkBuilder::empty(), &[]);
let empty_paths: Vec<&OsStr> = Vec::new();
assert_paths(td.path(), &WalkBuilder::from_iter(empty_paths), &[]);
}
#[test]
fn from_iter() {
let td = tmpdir();
mkdirp(td.path().join("a/b/c"));
mkdirp(td.path().join("d/e/f"));
mkdirp(td.path().join("x/y"));
wfile(td.path().join("a/b/foo"), "");
wfile(td.path().join("d/e/f/foo"), "");
wfile(td.path().join("x/y/foo"), "");
let paths = vec![
td.path().join("a"),
td.path().join("d"),
td.path().join("x"),
];
assert_paths(
td.path(),
&WalkBuilder::from_iter(paths),
&[
"x",
"x/y",
"x/y/foo",
"d",
"d/e",
"d/e/f",
"d/e/f/foo",
"a",
"a/b",
"a/b/foo",
"a/b/c",
],
);
}
// This should always panic and never hang.
//
// Ref: https://github.com/BurntSushi/ripgrep/issues/3009
#[test]
#[should_panic]
fn panic_in_parallel() {
let td = tmpdir();
wfile(td.path().join("foo.txt"), "");
WalkBuilder::new(td.path())
.threads(40)
.build_parallel()
.run(|| Box::new(|_| panic!("oops!")));
}
// This should always panic and never hang. The first call to the visitor
// builder is used while processing the root paths. Previously, a panic on
// the third call occurred after the first worker had already been spawned,
// leaving it waiting indefinitely for workers that were never created.
#[test]
#[should_panic(expected = "builder panic")]
fn panic_in_parallel_builder() {
let td = tmpdir();
wfile(td.path().join("foo.txt"), "");
let mut builds = 0;
WalkBuilder::new(td.path()).threads(2).build_parallel().run(|| {
builds += 1;
if builds == 3 {
panic!("builder panic");
}
Box::new(|_| WalkState::Continue)
});
}
}

View file

@ -2,7 +2,8 @@ use std::path::Path;
use ignore::gitignore::{Gitignore, GitignoreBuilder};
const IGNORE_FILE: &'static str = "tests/gitignore_matched_path_or_any_parents_tests.gitignore";
const IGNORE_FILE: &'static str =
"tests/gitignore_matched_path_or_any_parents_tests.gitignore";
fn get_gitignore() -> Gitignore {
let mut builder = GitignoreBuilder::new("ROOT");
@ -23,7 +24,9 @@ fn test_path_should_be_under_root() {
#[test]
fn test_files_in_root() {
let gitignore = get_gitignore();
let m = |path: &str| gitignore.matched_path_or_any_parents(Path::new(path), false);
let m = |path: &str| {
gitignore.matched_path_or_any_parents(Path::new(path), false)
};
// 0x
assert!(m("ROOT/file_root_00").is_ignore());
@ -53,7 +56,9 @@ fn test_files_in_root() {
#[test]
fn test_files_in_deep() {
let gitignore = get_gitignore();
let m = |path: &str| gitignore.matched_path_or_any_parents(Path::new(path), false);
let m = |path: &str| {
gitignore.matched_path_or_any_parents(Path::new(path), false)
};
// 0x
assert!(m("ROOT/parent_dir/file_deep_00").is_ignore());
@ -83,8 +88,9 @@ fn test_files_in_deep() {
#[test]
fn test_dirs_in_root() {
let gitignore = get_gitignore();
let m =
|path: &str, is_dir: bool| gitignore.matched_path_or_any_parents(Path::new(path), is_dir);
let m = |path: &str, is_dir: bool| {
gitignore.matched_path_or_any_parents(Path::new(path), is_dir)
};
// 00
assert!(m("ROOT/dir_root_00", true).is_ignore());
@ -186,20 +192,25 @@ fn test_dirs_in_root() {
#[test]
fn test_dirs_in_deep() {
let gitignore = get_gitignore();
let m =
|path: &str, is_dir: bool| gitignore.matched_path_or_any_parents(Path::new(path), is_dir);
let m = |path: &str, is_dir: bool| {
gitignore.matched_path_or_any_parents(Path::new(path), is_dir)
};
// 00
assert!(m("ROOT/parent_dir/dir_deep_00", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_00/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_00/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_00/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_00/child_dir/file", false).is_ignore()
);
// 01
assert!(m("ROOT/parent_dir/dir_deep_01", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_01/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_01/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_01/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_01/child_dir/file", false).is_ignore()
);
// 02
assert!(m("ROOT/parent_dir/dir_deep_02", true).is_none());
@ -241,51 +252,67 @@ fn test_dirs_in_deep() {
assert!(m("ROOT/parent_dir/dir_deep_20", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_20/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_20/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_20/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_20/child_dir/file", false).is_ignore()
);
// 21
assert!(m("ROOT/parent_dir/dir_deep_21", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_21/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_21/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_21/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_21/child_dir/file", false).is_ignore()
);
// 22
// dir itself doesn't match
assert!(m("ROOT/parent_dir/dir_deep_22", true).is_none());
assert!(m("ROOT/parent_dir/dir_deep_22/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_22/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_22/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_22/child_dir/file", false).is_ignore()
);
// 23
// dir itself doesn't match
assert!(m("ROOT/parent_dir/dir_deep_23", true).is_none());
assert!(m("ROOT/parent_dir/dir_deep_23/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_23/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_23/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_23/child_dir/file", false).is_ignore()
);
// 30
assert!(m("ROOT/parent_dir/dir_deep_30", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_30/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_30/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_30/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_30/child_dir/file", false).is_ignore()
);
// 31
assert!(m("ROOT/parent_dir/dir_deep_31", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_31/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_31/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_31/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_31/child_dir/file", false).is_ignore()
);
// 32
// dir itself doesn't match
assert!(m("ROOT/parent_dir/dir_deep_32", true).is_none());
assert!(m("ROOT/parent_dir/dir_deep_32/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_32/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_32/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_32/child_dir/file", false).is_ignore()
);
// 33
// dir itself doesn't match
assert!(m("ROOT/parent_dir/dir_deep_33", true).is_none());
assert!(m("ROOT/parent_dir/dir_deep_33/file", false).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_33/child_dir", true).is_ignore());
assert!(m("ROOT/parent_dir/dir_deep_33/child_dir/file", false).is_ignore());
assert!(
m("ROOT/parent_dir/dir_deep_33/child_dir/file", false).is_ignore()
);
}

View file

@ -20,6 +20,7 @@ ignore = { path = "../ignore" }
regex = "1.11.1"
[dev-dependencies]
insta = "1.48.0"
tempfile = "3.13.0"
pretty_assertions = "1.4.1"
unicode-width = "0.2.0"

View file

@ -48,7 +48,7 @@ static IGNORED_EXTENSIONS_GLOB: sync::LazyLock<String> = sync::LazyLock::new(||
)
});
pub static BINARY_EXTENSIONS_GLOB: sync::LazyLock<String> = sync::LazyLock::new(|| {
static BINARY_EXTENSIONS_GLOB: sync::LazyLock<String> = sync::LazyLock::new(|| {
format!(
"*.{{{}}}",
include_str!("fixtures/binary-extensions.txt")

View file

@ -2,5 +2,6 @@ package-lock.json
pnpm-lock.yaml
bun.lockb
.gitignore
.ignore
.env
.env.*

View file

@ -10,11 +10,13 @@ use crate::scanner::sources::{
public_source_entries_to_private_source_entries, PublicSourceEntry, SourceEntry, Sources,
};
use crate::GlobEntry;
use auto_source_detection::BINARY_EXTENSIONS_GLOB;
use bstr::ByteSlice;
use fast_glob::glob_match;
use fxhash::{FxHashMap, FxHashSet};
use ignore::{gitignore::GitignoreBuilder, WalkBuilder};
use ignore::{
gitignore::{Gitignore, GitignoreBuilder},
WalkBuilder,
};
use init_tracing::{init_tracing, SHOULD_TRACE};
use rayon::prelude::*;
use std::path::{Path, PathBuf};
@ -22,16 +24,71 @@ use std::sync::{Arc, Mutex};
use std::time::SystemTime;
use tracing::event;
// @source "some/folder"; // This is auto source detection
// @source "some/folder/**/*"; // This is auto source detection
// @source "some/folder/*.html"; // This is just a glob, but new files matching this should be included
// @source "node_modules/my-ui-lib"; // Auto source detection but since node_modules is explicit we allow it
// // Maybe could be considered `external(…)` automatically if:
// // 1. It's git ignored but listed explicitly
// // 2. It exists outside of the current working directory (do we know that?)
// # `@source` semantics
//
// @source "do-include-me.bin"; // `.bin` is typically ignored, but now it's explicit so should be included
// @source "git-ignored.html"; // A git ignored file that is listed explicitly, should be scanned
// Every `@source` directive is classified as one of:
//
// - `Auto`: `@source "some/folder"` or `@source "some/folder/**/*"` — auto source detection.
// The folder is scanned recursively while respecting `.gitignore` files and the default rules
// (skip `node_modules`/`.git`/…, skip binary and irrelevant extensions, skip lock files, …).
//
// - `External`: an `Auto` source whose folder is itself ignored (by a `.gitignore` or because
// it's a default-ignored directory like `node_modules`), e.g.
// `@source "node_modules/my-ui-lib"`. Since the folder was listed explicitly, its ignoredness
// is bypassed: everything inside is scanned as if it were an `Auto` source, except that
// `.gitignore` files from at or above the folder no longer apply — they (including the
// self-ignoring `*` file that generators typically place inside such folders) are what made
// it ignored in the first place. `.gitignore` files *deeper inside* the folder still apply,
// and so do the default rules: nested `node_modules`, binary extensions, etc. stay ignored.
//
// - `Pattern`: `@source "some/folder/*.html"` — an explicit glob. Only files matching the glob
// are scanned. The *static* prefix of the glob (`some/folder`) is the explicit part: it is
// reached even when it is git ignored or hidden behind a default-ignored directory
// (`@source "node_modules/lib/dist/*.html"` works). The *wildcard* part is not explicit:
// while expanding it we still respect `.gitignore` files inside the walked subtree and the
// default-ignored directories (`@source "./**/*.html"` does not descend into `node_modules`
// or a git ignored `dist/`; `@source "./dist/**/*.html"` does descend into `dist/`).
// Individual *files* matching the glob are always included, even when git ignored — you were
// explicit about wanting files of that shape (`@source "git-ignored.html"` and
// `@source "*.styl"` work). Files that are ignored by default are only included when the
// pattern is explicit about them: it names a concrete file (`@source "do-include-me.bin"`,
// `@source ".env"`) or pins an extension (`@source "logo.{jpg,png}"`). A glob that does
// neither (e.g. `@source "blog/*/post/**/*"`) still applies the default file rules.
//
// - `Ignored`: `@source not "…"` — excludes matching files/folders, even when a `.gitignore`
// or another `@source` allows them.
//
// Later directives win over earlier ones on conflict: `@source not "./x"` followed by
// `@source "./x/keep.html"` scans `keep.html`, and vice versa excludes it. The same holds for
// directory sources: `@source not "./x"` followed by `@source "./x"` re-includes `./x` (with
// the normal auto source detection rules applied inside).
//
// # Implementation
//
// All sources are scanned in a single file system walk. The vendored `ignore` crate only
// provides the traversal itself (parallel walking, symlink loop handling); its built-in
// gitignore handling is disabled, because it computes one global verdict per path while the
// `@source` semantics are per source: the same directory can be pruned for an auto source but
// walkable for a pattern source, and a file can be gitignored for an auto source but rescued
// by a glob.
//
// Instead, the [`Resolver`] implements the semantics in the walker's `filter_entry` callback,
// backed by its own lazily-loaded cache of the on-disk ignore files (`.gitignore`, `.ignore`,
// and the repository's `.git/info/exclude`, applied up to the git repository root):
//
// - The walk roots are the "maximal" source bases; nested bases are reached by walking, and
// the resolver keeps the static path towards an explicitly listed base open even through
// ignored directories.
//
// - A directory is entered when at least one source can contribute files inside of it, where
// each source kind applies its own rules: auto sources check the default rules and the full
// gitignore chain, external sources only the default rules, and pattern sources check the
// default rules, the `.gitignore` files at or below their base, and whether the glob can
// match anything inside the directory. Directories that cannot contribute anything are
// never descended into.
//
// - A file is kept when at least one source includes it, honoring the directive order of
// `@source not` (the later directive wins).
#[derive(Debug, Clone)]
pub enum ChangedContent {
@ -63,6 +120,9 @@ pub struct Scanner {
/// The walker to detect all files that we have to scan
walker: Option<WalkBuilder>,
/// The resolver implementing the `@source` semantics for the walker
resolver: Option<Arc<Resolver>>,
/// All found extensions
extensions: FxHashSet<String>,
@ -111,11 +171,13 @@ impl Scanner {
}
}
let walker = create_walker(&sources);
let resolver = Arc::new(Resolver::new(&sources));
let walker = create_walker(resolver.clone());
Self {
sources,
walker,
resolver: Some(resolver),
..Default::default()
}
}
@ -408,7 +470,16 @@ impl Scanner {
for entry in all_entries {
match entry {
WalkEntry::Dir(path) => {
self.dirs.insert(path);
// Directories that are only walked to reach an explicitly listed base are
// not part of any source's content: they must not widen the generated
// file watcher globs.
let contributes = self
.resolver
.as_ref()
.is_some_and(|resolver| resolver.contributes_dir(&path));
if contributes {
self.dirs.insert(path);
}
}
WalkEntry::File {
path,
@ -707,205 +778,639 @@ fn walk_parallel(walker: &mut WalkBuilder) -> Vec<WalkEntry> {
Arc::try_unwrap(collected).unwrap().into_inner().unwrap()
}
/// Sets up a WalkBuilder with all source roots, gitignore rules, and source pattern matching.
/// Sets up the single walker for all sources.
///
/// This is the common setup shared between the full walker (with mtime tracking for re-scans)
/// and the parallel walker (without mtime tracking for the initial scan).
fn create_walker(sources: &Sources) -> Option<WalkBuilder> {
let mut other_roots: FxHashSet<&PathBuf> = FxHashSet::default();
let mut first_root: Option<&PathBuf> = None;
/// The walker is only used for the (parallel, symlink aware) traversal itself: all ignore
/// semantics are implemented by the [`Resolver`] in the `filter_entry` callback. The crate's
/// built-in gitignore handling can't be used because it computes one global verdict per path,
/// while the `@source` semantics are per source (see the spec at the top of this file): the
/// same directory can be pruned for an auto source but walkable for a pattern source, and a
/// file can be gitignored for an auto source but rescued by a glob.
///
/// The walk roots are the "maximal" source bases: a base contained in another base is reached
/// by walking, which the resolver allows even through ignored directories (the static path to
/// an explicit base bypasses ignore rules).
fn create_walker(resolver: Arc<Resolver>) -> Option<WalkBuilder> {
let mut roots = resolver.walk_roots().into_iter();
let first_root = roots.next()?;
let mut ignores: Vec<(&PathBuf, Vec<String>)> = Default::default();
let mut emit = |base, pattern| match ignores.last_mut() {
Some((prev_base, patterns)) if *prev_base == base => {
patterns.push(pattern);
}
_ => {
ignores.push((base, vec![pattern]));
}
};
for source in sources.iter() {
match source {
SourceEntry::Auto { base } => {
if first_root.is_none() {
first_root = Some(base);
} else {
other_roots.insert(base);
}
}
SourceEntry::Pattern { base, pattern } => {
let pattern = pattern.to_owned();
if first_root.is_none() {
first_root = Some(base);
} else {
other_roots.insert(base);
}
if !pattern.contains("**") {
// Specific patterns should take precedence even over git-ignored files:
emit(base, format!("!{}", pattern));
} else {
// Assumption: the pattern we receive will already be brace expanded. So
// `*.{html,jsx}` will result in two separate patterns: `*.html` and `*.jsx`.
if let Some(extension) = Path::new(&pattern).extension() {
// Extend auto source detection to include the extension
emit(base, format!("!*.{}", extension.to_string_lossy()));
}
}
}
SourceEntry::Ignored { base, pattern } => {
emit(base, pattern.to_owned());
}
SourceEntry::External { base } => {
if first_root.is_none() {
first_root = Some(base);
} else {
other_roots.insert(base);
}
// External sources should take precedence even over git-ignored files:
emit(base, "!/**/*".to_owned());
// External sources should still disallow binary extensions:
emit(base, BINARY_EXTENSIONS_GLOB.clone());
}
}
let mut builder = WalkBuilder::new(first_root);
for root in roots {
builder.add(root);
}
let mut builder = WalkBuilder::new(first_root?);
// We have to follow symlinks
builder.follow_links(true);
// Scan hidden files / directories
builder.hidden(false);
// Disable all of the built-in filtering (hidden files, .gitignore files, parent
// directories, global gitignore files, …): the resolver implements the ignore semantics.
builder.standard_filters(false);
// Don't respect global gitignore files
builder.git_global(false);
// By default, allow .gitignore files to be used regardless of whether or not
// a .git directory is present. This is an optimization for when projects
// are first created and may not be in a git repo yet.
builder.require_git(false);
// If we are in a git repo then require it to ensure that only rules within
// the repo are used. For example, we don't want to consider a .gitignore file
// in the user's home folder if we're in a git repo.
//
// The alternative is using a call like `.parents(false)` but that will
// prevent looking at parent directories for .gitignore files from within
// the repo and that's not what we want.
//
// For example, in a project with this structure:
//
// home
// .gitignore
// my-project
// .gitignore
// apps
// .gitignore
// web
// {root}
//
// We do want to consider all .gitignore files listed:
// - home/.gitignore
// - my-project/.gitignore
// - my-project/apps/.gitignore
//
// However, if a repo is initialized inside my-project then only the following
// make sense for consideration:
// - my-project/.gitignore
// - my-project/apps/.gitignore
//
// Setting the require_git(true) flag conditionally allows us to do this.
for parent in first_root?.ancestors() {
if parent.join(".git").exists() {
builder.require_git(true);
break;
}
}
for root in other_roots {
builder.add(root);
}
// Setup auto source detection rules
for ignore in auto_source_detection::RULES.iter() {
builder.add_gitignore(ignore.clone());
}
// Setup ignores based on `@source` definitions
for (base, patterns) in ignores {
let mut ignore_builder = GitignoreBuilder::new(base);
for pattern in patterns {
ignore_builder.add_line(None, &pattern).unwrap();
}
let ignore = ignore_builder.build().unwrap();
builder.add_gitignore(ignore);
}
// Pre-compute source matching data to avoid allocations in the hot filter_entry path
let auto_bases: Vec<PathBuf> = sources
.iter()
.filter_map(|source| match source {
SourceEntry::Auto { base } | SourceEntry::External { base } => Some(base.clone()),
_ => None,
})
.collect();
let pattern_sources: Vec<(PathBuf, String)> = sources
.iter()
.filter_map(|source| match source {
SourceEntry::Pattern { base, pattern } => Some((base.into(), pattern.into())),
_ => None,
})
.collect();
// Source pattern matching filter (lock-free, safe for parallel walking)
builder.filter_entry(move |entry| {
let path = entry.path();
// Ensure the entries are matching any of the provided source patterns (this is
// necessary for manual-patterns that can filter the file extension)
if path.is_file() {
let mut matches = false;
for base in &auto_bases {
if path.starts_with(base) {
matches = true;
break;
}
}
if !matches {
for (base, pattern) in &pattern_sources {
let remainder = path.strip_prefix(base);
if remainder.is_ok_and(|remainder| {
let mut path_str = remainder.to_string_lossy().to_string();
if !path_str.starts_with("/") {
path_str = format!("/{path_str}");
}
glob_match(pattern, path_str.as_bytes())
}) {
matches = true;
break;
}
}
}
if !matches {
return false;
}
}
true
});
builder.filter_entry(move |entry| resolver.keep(entry));
Some(builder)
}
/// Implements the `@source` semantics for a single file system walk.
///
/// For every walked entry, the resolver computes the union of the per-source verdicts:
///
/// - a directory is entered when at least one source can contribute files inside of it
/// - a file is kept when at least one source includes it, honoring the directive order of
/// `@source not` (the later directive wins)
///
/// The per-source verdicts require gitignore decisions relative to different anchors (see
/// [`Boundary`]: each source kind respects a different part of the `.gitignore` chain), so the
/// resolver maintains its own lazily-loaded cache of ignore files instead of using the
/// walker's built-in handling.
#[derive(Debug)]
struct Resolver {
/// `Auto` sources: directive position and base
autos: Vec<(usize, PathBuf)>,
/// `External` sources: directive position and base
externals: Vec<(usize, PathBuf)>,
/// `Pattern` sources, grouped by base
patterns: Vec<PatternGroup>,
/// `@source not` directives
nots: Vec<NotRule>,
/// Lazily loaded ignore files (`.gitignore`, `.ignore`, `.git/info/exclude`) per directory
ignore_files: IgnoreFiles,
/// Memoized "is this directory reachable for an auto source": every directory on the path
/// from an auto base down to it passes the gitignore chain and the default rules
auto_reachable: Mutex<FxHashMap<PathBuf, bool>>,
/// Memoized "is this directory reachable for an external source": like `auto_reachable`,
/// but only `.gitignore` files strictly below the external base apply
external_reachable: Mutex<FxHashMap<PathBuf, bool>>,
}
/// All `Pattern` sources sharing a base, e.g. `@source "src/*.{html,jsx}"` produces the
/// patterns `/*.html` and `/*.jsx` for the base `src`. Patterns carry the position of their
/// `@source` directive to resolve conflicts with `@source not` directives: the later
/// directive wins.
#[derive(Debug)]
struct PatternGroup {
base: PathBuf,
/// The glob patterns, relative to `base`, with their directive positions
patterns: Vec<(usize, String)>,
/// Memoized "is this directory reachable for this group": like
/// `Resolver::auto_reachable`, but only `.gitignore` files at or below the base apply —
/// the static base is explicit, so everything above it is bypassed
reachable: Mutex<FxHashMap<PathBuf, bool>>,
}
impl Resolver {
fn new(sources: &Sources) -> Self {
let mut autos = vec![];
let mut externals = vec![];
let mut patterns: Vec<PatternGroup> = vec![];
let mut nots = vec![];
for (idx, source) in sources.iter().enumerate() {
match source {
SourceEntry::Auto { base } => autos.push((idx, base.clone())),
SourceEntry::External { base } => externals.push((idx, base.clone())),
SourceEntry::Pattern { base, pattern } => {
match patterns.iter_mut().find(|group| &group.base == base) {
Some(group) => group.patterns.push((idx, pattern.clone())),
None => patterns.push(PatternGroup {
base: base.clone(),
patterns: vec![(idx, pattern.clone())],
reachable: Default::default(),
}),
}
}
SourceEntry::Ignored { base, pattern } => nots.push(NotRule {
idx,
base: base.clone(),
pattern: pattern.clone(),
}),
}
}
Self {
autos,
externals,
patterns,
nots,
ignore_files: IgnoreFiles::default(),
auto_reachable: Default::default(),
external_reachable: Default::default(),
}
}
/// All source bases
fn bases(&self) -> impl Iterator<Item = &PathBuf> {
self.autos
.iter()
.map(|(_, base)| base)
.chain(self.externals.iter().map(|(_, base)| base))
.chain(self.patterns.iter().map(|group| &group.base))
}
/// The walk roots: all bases that are not contained in another base. Nested bases are
/// reached by walking (the resolver keeps the path to an explicit base open).
fn walk_roots(&self) -> Vec<PathBuf> {
let mut roots: Vec<PathBuf> = vec![];
for base in self.bases() {
if self
.bases()
.any(|other| other != base && base.starts_with(other))
{
continue;
}
if !roots.contains(base) {
roots.push(base.clone());
}
}
roots
}
/// Whether to keep the given walk entry
fn keep(&self, entry: &ignore::DirEntry) -> bool {
// Always keep the walk roots themselves; they are explicitly listed bases
if entry.depth() == 0 {
return true;
}
let path = entry.path();
let is_dir = entry.file_type().map(|ft| ft.is_dir()).unwrap_or(false);
if is_dir {
self.keep_dir(path)
} else {
self.keep_file(path)
}
}
/// Whether to keep walking the given directory: either a source can contribute files
/// inside of it, or it is on the static path towards an explicitly listed base.
fn keep_dir(&self, dir: &Path) -> bool {
// A directory on the static path to an explicit base can always be entered (the base
// is explicit, so its ignoredness is bypassed), unless everything below it is excluded
// again by a later `@source not` directive.
let leads_to_base = |idx: usize, base: &PathBuf| {
base != dir
&& base.starts_with(dir)
&& !self
.nots
.iter()
.any(|not| not.idx > idx && not.matches(base))
};
if self.autos.iter().any(|(idx, base)| leads_to_base(*idx, base))
|| self
.externals
.iter()
.any(|(idx, base)| leads_to_base(*idx, base))
|| self.patterns.iter().any(|group| {
group
.patterns
.iter()
.any(|(idx, _)| leads_to_base(*idx, &group.base))
})
{
return true;
}
self.contributes_dir(dir)
}
/// Whether at least one source can contribute files inside the given directory. Unlike
/// [`Resolver::keep_dir`], directories that are only walked to reach an explicitly listed
/// base don't count: they are not part of any source's content, e.g. for the purpose of
/// generating file watcher globs.
fn contributes_dir(&self, dir: &Path) -> bool {
// Some source must be able to contribute files inside the directory, and not be
// overridden by a later `@source not` directive.
let not_after = |idx: usize| {
self.nots
.iter()
.any(|not| not.idx > idx && not.matches(dir))
};
if self
.autos
.iter()
.any(|(idx, _)| self.auto_reachable(dir) && !not_after(*idx))
{
return true;
}
if self
.externals
.iter()
.any(|(idx, _)| self.external_reachable(dir) && !not_after(*idx))
{
return true;
}
self.patterns.iter().any(|group| {
dir.strip_prefix(&group.base).is_ok_and(|remainder| {
group.patterns.iter().any(|(idx, pattern)| {
dir_could_contain_matches(pattern, remainder) && !not_after(*idx)
}) && self.pattern_reachable(group, dir)
})
})
}
/// Whether at least one source includes the given file
fn keep_file(&self, file: &Path) -> bool {
let Some(parent) = file.parent() else {
return false;
};
let not_after = |idx: usize| {
self.nots
.iter()
.any(|not| not.idx > idx && not.matches(file))
};
// Auto sources: the file must pass the default rules and the gitignore chain
if self.autos.iter().any(|(idx, base)| {
file.starts_with(base)
&& self.auto_reachable(parent)
&& !is_ignored_by_default_rules(file, false)
&& !self
.ignore_files
.is_ignored(file, false, parent, Boundary::None)
&& !not_after(*idx)
}) {
return true;
}
// External sources: like auto sources, but only `.gitignore` files strictly below the
// base apply. Rules from at or above the base are bypassed — they are what made the
// directory ignored, and it was listed explicitly anyway — while `.gitignore` files
// deeper inside the external tree still apply.
if self.externals.iter().any(|(idx, base)| {
file.starts_with(base)
&& self.external_reachable(parent)
&& !is_ignored_by_default_rules(file, false)
&& !self
.ignore_files
.is_ignored(file, false, parent, Boundary::Inside(base))
&& !not_after(*idx)
}) {
return true;
}
// Pattern sources: the file must match a glob. A match beats file-level gitignore
// rules — you were explicit about wanting files of that shape — and the default file
// rules only apply when the pattern isn't explicit about the file's shape.
self.patterns.iter().any(|group| {
file.strip_prefix(&group.base).is_ok_and(|remainder| {
let remainder = rooted_posix(remainder);
group.patterns.iter().any(|(idx, pattern)| {
glob_match(pattern, remainder.as_bytes())
&& (pattern_bypasses_default_file_rules(pattern)
|| !is_ignored_by_default_rules(file, false))
&& !not_after(*idx)
}) && self.pattern_reachable(group, parent)
})
})
}
/// Whether the given directory is reachable for an auto source: every directory on the
/// path from the auto base down to it passes the default rules and the gitignore chain.
fn auto_reachable(&self, dir: &Path) -> bool {
reachable(
&self.auto_reachable,
dir,
|dir| self.autos.iter().any(|(_, base)| base == dir),
|dir| {
!is_ignored_by_default_rules(dir, true)
&& dir.parent().is_some_and(|parent| {
!self.ignore_files.is_ignored(dir, true, parent, Boundary::None)
})
},
)
}
/// Like [`Resolver::auto_reachable`], but for external sources: only `.gitignore` files
/// strictly below the external base apply (see [`Boundary`]). Rules from at or above the
/// base are bypassed — they are what made the directory ignored, and it was listed
/// explicitly anyway.
///
/// When external bases are nested, the deepest base containing the directory bounds the
/// chain: reachability from an outer base implies reachability from a nested base (the
/// outer chain checks a superset of the ignore files), so this computes the union of the
/// per-base verdicts.
fn external_reachable(&self, dir: &Path) -> bool {
reachable(
&self.external_reachable,
dir,
|dir| self.externals.iter().any(|(_, base)| base == dir),
|dir| {
!is_ignored_by_default_rules(dir, true)
&& dir.parent().is_some_and(|parent| {
let boundary = self
.externals
.iter()
.map(|(_, base)| base)
.filter(|base| dir.starts_with(base))
.max_by_key(|base| base.components().count())
.map_or(Boundary::None, |base| Boundary::Inside(base));
!self.ignore_files.is_ignored(dir, true, parent, boundary)
})
},
)
}
/// Like [`Resolver::auto_reachable`], but for a pattern group: only `.gitignore` files at
/// or below the base apply — the static base is explicit, so everything above it is
/// bypassed.
fn pattern_reachable(&self, group: &PatternGroup, dir: &Path) -> bool {
reachable(
&group.reachable,
dir,
|dir| dir == group.base,
|dir| {
!is_ignored_by_default_rules(dir, true)
&& dir.parent().is_some_and(|parent| {
!self
.ignore_files
.is_ignored(dir, true, parent, Boundary::At(&group.base))
})
},
)
}
}
/// Whether every directory on the path from a source base down to `dir` passes the source's
/// `enter` rule. Bases themselves are always reachable: they are explicitly listed (and an
/// auto base that is itself ignored would have been promoted to an external source). Verdicts
/// are memoized per directory.
fn reachable(
memo: &Mutex<FxHashMap<PathBuf, bool>>,
dir: &Path,
is_base: impl Fn(&Path) -> bool,
enter: impl Fn(&Path) -> bool,
) -> bool {
// Walk up to the nearest base or directory with a memoized verdict…
let mut pending = vec![];
let mut current = dir;
let mut reachable = loop {
if is_base(current) {
break true;
}
if let Some(reachable) = memo.lock().unwrap().get(current) {
break *reachable;
}
pending.push(current.to_path_buf());
match current.parent() {
Some(parent) => current = parent,
// Reached the file system root without finding a base
None => break false,
}
};
// …then fill in the verdicts back down towards `dir`
for dir in pending.into_iter().rev() {
reachable = reachable && enter(&dir);
memo.lock().unwrap().insert(dir, reachable);
}
reachable
}
/// A lazily-loaded cache of the on-disk ignore files (`.gitignore`, `.ignore` and the
/// repository's `.git/info/exclude`).
#[derive(Debug, Default)]
struct IgnoreFiles {
/// The matcher chains per directory: all matchers that apply to paths inside the
/// directory, deepest first, from the directory itself up to the repository root (or the
/// file system root outside of a git repository, matching `git init`-less projects where
/// all ancestor `.gitignore` files apply)
chains: Mutex<FxHashMap<PathBuf, Arc<Vec<Arc<Gitignore>>>>>,
}
/// Which part of the ignore file chain applies to a source, anchored at its base:
///
/// - `Auto` sources respect the full chain, up to the git repository root
/// - `Pattern` sources respect ignore files at or below their base (the static base is
/// explicit, everything above it is bypassed)
/// - `External` sources respect ignore files strictly below their base: the base's own
/// ignore file is part of its bypassed ignoredness — inside ignored trees it is typically
/// the self-ignoring `*` file that generators drop into the directory — while deeper ignore
/// files are deliberate signals about specific contents and still apply
#[derive(Debug, Clone, Copy)]
enum Boundary<'a> {
None,
At(&'a Path),
Inside(&'a Path),
}
impl Boundary<'_> {
/// Whether an ignore file rooted at the given directory applies
fn applies_to(&self, dir: &Path) -> bool {
match self {
Boundary::None => true,
Boundary::At(base) => dir.starts_with(base),
Boundary::Inside(base) => dir != *base && dir.starts_with(base),
}
}
}
impl IgnoreFiles {
/// Whether the ignore files definitively ignore the given path. `dir` is the directory
/// containing the path, and `boundary` restricts which ignore files of the chain apply.
///
/// The deepest ignore file with a definitive answer wins, matching git's precedence, so a
/// path that a deeper ignore file re-includes via a `!` pattern is not ignored.
fn is_ignored(&self, path: &Path, is_dir: bool, dir: &Path, boundary: Boundary) -> bool {
for matcher in self.chain(dir).iter() {
if !boundary.applies_to(matcher.path()) {
// Chains are ordered deepest first, so nothing below the boundary can follow
break;
}
match matcher.matched(path, is_dir) {
ignore::Match::Ignore(_) => return true,
ignore::Match::Whitelist(_) => return false,
ignore::Match::None => {}
}
}
false
}
/// The matcher chain for paths inside the given directory: the directory's own matcher
/// first, then its parents' matchers, up to and including the git repository root.
fn chain(&self, dir: &Path) -> Arc<Vec<Arc<Gitignore>>> {
if let Some(chain) = self.chains.lock().unwrap().get(dir) {
return chain.clone();
}
let is_repo_root = dir.join(".git").exists();
let mut chain = vec![];
if let Some(matcher) = load_ignore_files(dir, is_repo_root) {
chain.push(matcher);
}
// Stop at the git repository root so that ignore files outside of the repository are
// not considered. Without a repository, all ancestor ignore files apply.
if !is_repo_root {
if let Some(parent) = dir.parent() {
chain.extend(self.chain(parent).iter().cloned());
}
}
let chain = Arc::new(chain);
self.chains
.lock()
.unwrap()
.insert(dir.to_path_buf(), chain.clone());
chain
}
}
/// Compile the ignore rules of the given directory, combining (from low to high precedence)
/// the repository's `.git/info/exclude`, the `.gitignore` file, and the `.ignore` file.
fn load_ignore_files(dir: &Path, is_repo_root: bool) -> Option<Arc<Gitignore>> {
let mut builder = GitignoreBuilder::new(dir);
let mut any = false;
let mut add = |file: PathBuf| {
if file.is_file() {
// I/O errors and partially invalid ignore files are ignored, matching the
// walker's behavior.
let _ = builder.add(file);
any = true;
}
};
if is_repo_root {
add(dir.join(".git").join("info").join("exclude"));
}
add(dir.join(".gitignore"));
add(dir.join(".ignore"));
if any {
builder.build().ok().map(Arc::new)
} else {
None
}
}
/// Whether a path is ignored by the default auto source detection rules: directories like
/// `node_modules`, binary and irrelevant extensions, lock files, … For directories only the
/// directory's own name is checked; the path towards it is checked by the reachability
/// helpers one directory at a time.
fn is_ignored_by_default_rules(path: &Path, is_dir: bool) -> bool {
auto_source_detection::RULES
.iter()
.any(|ignore| ignore.matched(path, is_dir).is_ignore())
}
/// An `@source not` directive, together with its position.
#[derive(Debug, Clone)]
struct NotRule {
/// Position of the `@source not` directive
idx: usize,
base: PathBuf,
/// The glob pattern, relative to `base`, e.g. `/ignored/**/*`
pattern: String,
}
impl NotRule {
/// Whether this directive excludes the given path: a file, or a directory and thereby
/// everything inside of it.
///
/// Like a gitignore rule, the pattern excludes a whole subtree when it matches a
/// directory, so besides the path itself every ancestor directory (up to the directive's
/// base) is tested as well. E.g. `@source not "./src/ba*"` excludes `src/bar/index.html`
/// because `/ba*` matches the `src/bar` directory. Note that directory-shaped directives
/// (`@source not "./some/dir"`) are normalized to a `/**/*` pattern with the directory as
/// its base, which matches everything inside the directory directly.
fn matches(&self, path: &Path) -> bool {
let Ok(remainder) = path.strip_prefix(&self.base) else {
return false;
};
// A directory-shaped directive (normalized to a `/**/*` pattern) also excludes the
// base directory itself, not just its contents, so the directory can be pruned.
if remainder.as_os_str().is_empty() {
return self.pattern == "/**/*";
}
remainder.ancestors().any(|prefix| {
!prefix.as_os_str().is_empty()
&& glob_match(&self.pattern, rooted_posix(prefix).as_bytes())
})
}
}
/// Serialize a path relative to some base as a `/`-rooted posix style string, e.g.
/// `/ba*/index.html`, matching how source patterns are stored.
fn rooted_posix(path: &Path) -> String {
let posix = crate::scanner::sources::path_to_posix_string(path);
if posix.starts_with('/') {
posix
} else {
format!("/{posix}")
}
}
/// Whether a directory (relative to the pattern's base) can contain files matching the pattern.
/// Used to prune directories that can never contribute, e.g. for `/ba*/*.html` only `ba*`
/// directories are entered.
fn dir_could_contain_matches(pattern: &str, dir: &Path) -> bool {
let pattern_components: Vec<&str> = pattern
.trim_start_matches('/')
.split('/')
.filter(|c| !c.is_empty())
.collect();
for (i, component) in dir.components().enumerate() {
let component = component.as_os_str().to_string_lossy();
// Once we see a `**` everything nested can contain matches
match pattern_components.get(i) {
Some(&"**") => return true,
// The last pattern component matches files, not directories. A directory nested
// deeper than the pattern's directory part can never contain matches.
Some(_) if i + 1 >= pattern_components.len() => return false,
Some(pattern_component) => {
if !glob_match(pattern_component, component.as_bytes()) {
return false;
}
}
None => return false,
}
}
true
}
/// Whether a pattern is explicit enough to bypass the default file rules.
///
/// A pattern without any wildcards names a concrete file, e.g. `/.env` or
/// `/do-include-me.bin` — you asked for exactly this file, so the default rules never apply.
/// A pattern that pins a specific extension, e.g. `/*.html` or `/**/*.bin`, bypasses them as
/// well. Patterns that do neither (e.g. `/blog/*/**/*`) keep the default file rules applied.
fn pattern_bypasses_default_file_rules(pattern: &str) -> bool {
// Concrete file, no wildcards (braces have already been expanded away)
if !pattern.contains(['*', '?', '[']) {
return true;
}
// Pinned extension
match Path::new(pattern).extension().and_then(|ext| ext.to_str()) {
Some(ext) => !ext.contains(['*', '?', '[']),
None => false,
}
}
#[cfg(test)]
mod tests {
use super::{ChangedContent, Scanner};

View file

@ -1,6 +1,5 @@
use crate::GlobEntry;
use bexpand::Expression;
use fxhash::{FxHashMap, FxHashSet};
use fxhash::FxHashMap;
use ignore::gitignore::Gitignore;
use std::path::{Component, Path, PathBuf};
use tracing::{event, Level};
@ -31,6 +30,24 @@ pub enum SourceEntry {
/// ```
Auto { base: PathBuf },
/// An `Auto` source whose directory is itself ignored (by the default rules, e.g.
/// `node_modules`, or by a `.gitignore`) but was explicitly listed anyway.
///
/// Represented by:
///
/// ```css
/// @source "../node_modules/my-lib";`
/// @source "../node_modules/my-lib/**/*";`
/// ```
///
/// Being explicit bypasses the ignoredness of the directory: everything inside is scanned
/// as if it were a regular auto source, except that `.gitignore` files from at or above
/// the directory no longer apply — they (including the self-ignoring `*` file that
/// generators typically place inside such directories) are what made it ignored in the
/// first place. `.gitignore` files *deeper inside* the directory still apply, and so do
/// the default rules, so e.g. nested `node_modules` stay ignored.
External { base: PathBuf },
/// Explicit source pattern regardless of any auto source detection rules
///
/// Represented by:
@ -48,18 +65,11 @@ pub enum SourceEntry {
/// @source not "src";`
/// @source not "src/**/*.html";`
/// ```
///
/// Note that directory-shaped directives (`@source not "src"`) are normalized to
/// `base: "src", pattern: "/**/*"`, which is semantically identical: everything under the
/// directory is ignored.
Ignored { base: PathBuf, pattern: String },
/// External sources are directories that are ignored (by us or .gitignore rules), but should be
/// included bypassing the default ignore rules.
///
/// Represented by:
///
/// ```css
/// @source "../node_modules/my-lib";`
/// @source "../node_modules/my-lib/**/*";`
/// ```
External { base: PathBuf },
}
#[derive(Debug, Clone, Default)]
@ -77,164 +87,6 @@ impl Sources {
}
}
/// When dealing with a pattern, then it could be that we end up with:
///
/// ```json
/// { base: '/some/folder', pattern: 'foo.ts' }
/// ```
///
/// If we just emit `!foo.ts` for the `/some/folder` path, then _everything_ else in that folder
/// would still be walked (but the result will be ignored).
///
/// Instead, we have to ensure that we ignore everything in that folder _except_ for the `foo.ts`
/// pattern.
///
/// This should be equivalent to:
/// ```gitignore
/// *
/// !foo.ts
/// ```
///
/// However, we have to be careful that we don't start ignoring files/folders that already exist.
/// ```css
/// @source "./some/folder/foo.ts";
/// @source "./some/folder/bar.ts";
/// ```
/// Would result in:
/// ```json
/// { base: '/some/folder', pattern: 'foo.ts' }
/// { base: '/some/folder', pattern: 'bar.ts' }
/// ```
///
/// If we were to blindly emit `*` for each pattern, then the `.gitignore` equivalent would look like
/// this:
/// ```gitignore
/// *
/// !foo.ts
/// *
/// !bar.ts
/// ```
///
/// This would result in ignoring the `foo.ts` file as well. Therefore we only want to insert
/// this `*` pattern when nothing else exists yet.
///
/// There is another problem that we need to solve. Let's say you have a pattern that contains a `*`
/// in the pattern:
/// ```css
/// @source './src/ba*/*.html';
/// ```
///
/// This would result in
/// ```json
/// { base: '/src', pattern: '/ba*/*.html' }
/// ```
///
/// If we now inject the `*` pattern for the `/src` folder, then we wouldn't scan any `ba*` folders
/// (e.g. `bar` or `baz`). For this, we have to make sure that we add inverse patterns for these
/// folders. This would essentially result in:
/// ```gitignore
/// * ← ignore everything
/// !/ba*/ ← except for the `ba*/` pattern, so we scan these folders
/// !/ba*/*.html ← then ensure we scan the `*.html` files in it as well
/// ```
///
fn expand_restricted_patterns(sources: Vec<SourceEntry>) -> Vec<SourceEntry> {
let unrestricted_roots = sources
.iter()
.filter_map(|source| match source {
SourceEntry::Auto { base } | SourceEntry::External { base } => Some(base.clone()),
SourceEntry::Pattern { base, pattern } if pattern.contains("**") => Some(base.clone()),
_ => None,
})
.collect::<Vec<_>>();
// Bases of restricted patterns. Each of these becomes its own walk root with its own
// `*` + `!<pattern>` rules, so an ancestor base must not ignore them recursively.
let pattern_roots = sources
.iter()
.filter_map(|source| match source {
SourceEntry::Pattern { base, .. } => Some(base.clone()),
_ => None,
})
.collect::<Vec<_>>();
let mut restricted_roots: FxHashSet<PathBuf> = FxHashSet::default();
let mut expanded = vec![];
for source in sources {
let SourceEntry::Pattern { base, pattern } = &source else {
expanded.push(source);
continue;
};
// `base` is already included by another `@source` that we know should be walked. This
// includes the case where `base` is _nested_ inside such a root, because everything under
// an unrestricted root is walked already. Restricting it would incorrectly hide siblings
// that the broader source is supposed to pick up.
if unrestricted_roots.iter().any(|root| base.starts_with(root)) {
expanded.push(source);
continue;
}
// Ignore everything in the directory. We will later add the specific patterns we are
// interested in.
if restricted_roots.insert(base.clone()) {
// When another source root is nested inside this base — an unrestricted root, or the
// base of another restricted pattern (which is walked from its own root with its own
// rules) — only ignore direct children so the nested root can still be walked.
let has_nested_root = unrestricted_roots.iter().any(|root| root.starts_with(base))
|| pattern_roots
.iter()
.any(|root| root != base && root.starts_with(base));
let pattern = if has_nested_root { "/*" } else { "*" };
expanded.push(SourceEntry::Ignored {
base: base.clone(),
pattern: pattern.to_owned(),
});
}
// Ensure to _include_ parent paths, otherwise the `*` from above would block walking the
// folders that need to be walked.
//
// ```css
// @source './src/ba*/*.html';
// ```
//
// ```gitignore
// * ← added by the above rule
// !/ba*/ ← this is what we're focusing on in this block
// !/ba*/*.html ← this is added later
// ```
{
let mut dir = PathBuf::new();
let mut components = Path::new(pattern).components().peekable();
while let Some(component) = components.next() {
if components.peek().is_none() {
break;
}
match component {
Component::Prefix(_) | Component::RootDir | Component::CurDir => continue,
Component::ParentDir | Component::Normal(_) => dir.push(component),
}
expanded.push(SourceEntry::Ignored {
base: base.clone(),
pattern: format!("!/{}/", path_to_posix_string(&dir).trim_start_matches('/')),
});
}
}
// Track the original source
expanded.push(source);
}
expanded
}
impl PublicSourceEntry {
/// Optimize the PublicSourceEntry by trying to move all the static parts of the pattern to the
/// base of the PublicSourceEntry.
@ -340,10 +192,16 @@ impl PublicSourceEntry {
else if !self.pattern.starts_with("/") {
self.pattern = format!("/{}", self.pattern);
}
// `src/**` means everything underneath `src`, just like `src/**/*` and `src` do.
// Normalize it so all three are classified as auto source detection.
if self.pattern == "/**" {
self.pattern = "/**/*".to_owned();
}
}
}
fn path_to_posix_string(path: &Path) -> String {
pub(crate) fn path_to_posix_string(path: &Path) -> String {
let mut parts = Vec::new();
let mut is_rooted = false;
@ -474,65 +332,32 @@ mod tests {
}
#[test]
fn concrete_patterns_are_expanded_to_restrict_their_base() {
fn optimize_normalizes_double_star_to_auto_source_detection() {
let dir = tempdir().unwrap();
fs::create_dir_all(dir.path().join("src")).unwrap();
let base = dunce::canonicalize(dir.path().join("src")).unwrap();
let mut source = PublicSourceEntry {
base: dir.path().to_string_lossy().to_string(),
pattern: "src/**".to_string(),
negated: false,
};
source.optimize();
assert_eq!(source.pattern, "/**/*");
// …and therefore `src/**` is classified as an auto source, like `src/**/*` and `src`
let base = dunce::canonicalize(dir.path().join("src")).unwrap();
let sources = public_source_entries_to_private_source_entries(vec![PublicSourceEntry {
base: dir.path().to_string_lossy().to_string(),
pattern: "src/foo.html".to_string(),
pattern: "src/**".to_string(),
negated: false,
}]);
assert_eq!(
sources,
vec![
SourceEntry::Ignored {
base: base.clone(),
pattern: "*".to_string(),
},
SourceEntry::Pattern {
base,
pattern: "/foo.html".to_string(),
},
]
);
assert_eq!(sources, vec![SourceEntry::Auto { base }]);
}
#[test]
fn restricted_patterns_include_parent_directory_allow_rules() {
let dir = tempdir().unwrap();
fs::create_dir_all(dir.path().join("src")).unwrap();
let base = dunce::canonicalize(dir.path().join("src")).unwrap();
let sources = public_source_entries_to_private_source_entries(vec![PublicSourceEntry {
base: dir.path().to_string_lossy().to_string(),
pattern: "src/ef*/*.html".to_string(),
negated: false,
}]);
assert_eq!(
sources,
vec![
SourceEntry::Ignored {
base: base.clone(),
pattern: "*".to_string(),
},
SourceEntry::Ignored {
base: base.clone(),
pattern: "!/ef*/".to_string(),
},
SourceEntry::Pattern {
base,
pattern: "/ef*/*.html".to_string(),
},
]
);
}
#[test]
fn unrestricted_sources_do_not_expand_patterns_for_the_same_base() {
fn sources_are_converted_in_order() {
let dir = tempdir().unwrap();
fs::create_dir_all(dir.path().join("src")).unwrap();
let base = dunce::canonicalize(dir.path().join("src")).unwrap();
@ -548,72 +373,6 @@ mod tests {
pattern: "src/foo.html".to_string(),
negated: false,
},
]);
assert_eq!(
sources,
vec![
SourceEntry::Auto { base: base.clone() },
SourceEntry::Pattern {
base,
pattern: "/foo.html".to_string(),
},
]
);
}
#[test]
fn restricted_parent_bases_do_not_open_unrelated_siblings() {
let dir = tempdir().unwrap();
let project = dir.path().join("Users").join("robin").join("docus-test");
fs::create_dir_all(&project).unwrap();
let users = dunce::canonicalize(dir.path().join("Users")).unwrap();
let project = dunce::canonicalize(project).unwrap();
let sources = public_source_entries_to_private_source_entries(vec![
PublicSourceEntry {
base: project.to_string_lossy().to_string(),
pattern: "**/*".to_string(),
negated: false,
},
PublicSourceEntry {
base: project.to_string_lossy().to_string(),
pattern: "../../app.config.ts".to_string(),
negated: false,
},
]);
assert_eq!(
sources,
vec![
SourceEntry::Auto {
base: project.clone(),
},
SourceEntry::Ignored {
base: users.clone(),
pattern: "/*".to_string(),
},
SourceEntry::Pattern {
base: users,
pattern: "/app.config.ts".to_string(),
},
]
);
}
#[test]
fn restricted_patterns_preserve_source_order() {
let dir = tempdir().unwrap();
fs::create_dir_all(dir.path().join("src")).unwrap();
let base = dunce::canonicalize(dir.path().join("src")).unwrap();
let sources = public_source_entries_to_private_source_entries(vec![
PublicSourceEntry {
base: dir.path().to_string_lossy().to_string(),
pattern: "src/foo.html".to_string(),
negated: false,
},
PublicSourceEntry {
base: dir.path().to_string_lossy().to_string(),
pattern: "src/foo.html".to_string(),
@ -624,10 +383,7 @@ mod tests {
assert_eq!(
sources,
vec![
SourceEntry::Ignored {
base: base.clone(),
pattern: "*".to_string(),
},
SourceEntry::Auto { base: base.clone() },
SourceEntry::Pattern {
base: base.clone(),
pattern: "/foo.html".to_string(),
@ -660,7 +416,10 @@ mod tests {
fs::create_dir_all(&base).unwrap();
let base = dunce::canonicalize(&base).unwrap();
assert_eq!(auto_source_entry(&base), SourceEntry::Auto { base });
assert_eq!(
auto_source_entry(&base),
SourceEntry::Auto { base }
);
}
#[test]
@ -670,7 +429,10 @@ mod tests {
fs::create_dir_all(&base).unwrap();
let base = dunce::canonicalize(&base).unwrap();
assert_eq!(auto_source_entry(&base), SourceEntry::External { base });
assert_eq!(
auto_source_entry(&base),
SourceEntry::External { base }
);
}
#[test]
@ -684,7 +446,10 @@ mod tests {
fs::create_dir_all(&base).unwrap();
let base = dunce::canonicalize(&base).unwrap();
assert_eq!(auto_source_entry(&base), SourceEntry::External { base });
assert_eq!(
auto_source_entry(&base),
SourceEntry::External { base }
);
}
#[test]
@ -698,7 +463,54 @@ mod tests {
fs::create_dir_all(&base).unwrap();
let base = dunce::canonicalize(&base).unwrap();
assert_eq!(auto_source_entry(&base), SourceEntry::External { base });
assert_eq!(
auto_source_entry(&base),
SourceEntry::External { base }
);
}
#[test]
fn folders_reincluded_by_a_deeper_gitignore_stay_auto_sources() {
let dir = tempdir().unwrap();
fs::create_dir_all(dir.path().join(".git")).unwrap();
fs::write(dir.path().join(".gitignore"), "generated/\n").unwrap();
// The deepest `.gitignore` with a definitive answer wins: the re-include is reachable
// (no parent directory of `generated` is excluded), so `generated` is not ignored.
fs::create_dir_all(dir.path().join("packages").join("app")).unwrap();
fs::write(
dir.path().join("packages").join("app").join(".gitignore"),
"!generated/\n",
)
.unwrap();
let base = dir.path().join("packages").join("app").join("generated");
fs::create_dir_all(&base).unwrap();
let base = dunce::canonicalize(&base).unwrap();
assert_eq!(
auto_source_entry(&base),
SourceEntry::Auto { base }
);
}
#[test]
fn folders_inside_excluded_directories_become_external_sources() {
let dir = tempdir().unwrap();
fs::create_dir_all(dir.path().join(".git")).unwrap();
fs::write(dir.path().join(".gitignore"), "parent/\n").unwrap();
// This whitelist is unreachable: `parent` itself is excluded, so git never descends
// into it and the re-include of `child` has no effect.
fs::create_dir_all(dir.path().join("parent")).unwrap();
fs::write(dir.path().join("parent").join(".gitignore"), "!child/\n").unwrap();
let base = dir.path().join("parent").join("child");
fs::create_dir_all(&base).unwrap();
let base = dunce::canonicalize(&base).unwrap();
assert_eq!(
auto_source_entry(&base),
SourceEntry::External { base }
);
}
#[test]
@ -711,7 +523,10 @@ mod tests {
fs::create_dir_all(&base).unwrap();
let base = dunce::canonicalize(&base).unwrap();
assert_eq!(auto_source_entry(&base), SourceEntry::Auto { base });
assert_eq!(
auto_source_entry(&base),
SourceEntry::Auto { base }
);
}
}
@ -778,74 +593,79 @@ pub fn public_source_entries_to_private_source_entries(
.map(|public_source| {
let mut source: SourceEntry = public_source.into();
// Promote auto-sources to external sources if they were gitignored
// Mark auto sources as external if their directory is gitignored
if let SourceEntry::Auto { ref base } = source {
let inside_git_repo = base.ancestors().any(|dir| dir.join(".git").exists());
// Walk up from the folder, applying each `.gitignore` relative to the directory
// that contains it (matching git), and stop at the git repository root so
// `.gitignore` files outside of the repo are not considered.
// The chain of directories whose `.gitignore` files can apply: `base` itself and
// its ancestors, up to and including the git repository root so `.gitignore`
// files outside of the repo are not considered.
//
// Without a git repository there is no repository root to stop at. Stop once the
// directory contains the current working directory instead, so `.gitignore`
// files outside of the project (e.g. in the user's home directory) can never
// promote a source to an external source. Note that the file walker still
// applies those `.gitignore` files when deciding which files to scan.
let mut chain: Vec<&Path> = vec![];
for dir in base.ancestors() {
let gitignore = gitignores.entry(dir.to_path_buf()).or_insert_with(|| {
let path = dir.join(".gitignore");
chain.push(dir);
// `Gitignore::new` roots the matcher at the directory containing the file,
// so patterns match relative to it.
path.is_file().then(|| Gitignore::new(&path).0)
});
// Only `.gitignore` files in ancestors of `base` can ignore `base` itself.
// Patterns in `base`'s own `.gitignore` only match paths _inside_ `base`, never
// `base` itself (the file walker still applies them to `base`'s contents).
//
// Skipping `base`'s own `.gitignore` also prevents a false positive for
// whitelist style `.gitignore` files, because relativizing `base` against
// itself yields the empty path, which incorrectly matches `/*`.
//
// E.g.:
//
// ```gitignore
// /*
// !/.gitignore
// !/app
// !/public
// ```
//
// Everything inside `base` except `.gitignore`, `app` and `public` is ignored,
// but `base` itself is not.
if dir != base {
if let Some(gitignore) = gitignore {
if gitignore
.matched_path_or_any_parents(&base, true)
.is_ignore()
{
source = SourceEntry::External { base: base.into() };
break;
}
}
}
// Stop at the git repository root.
if dir.join(".git").exists() {
break;
}
// Without a git repository there is no repository root to stop at. Stop once
// the directory contains the current working directory instead, so `.gitignore`
// files outside of the project (e.g. in the user's home directory) can never
// promote a source to an external source. Note that the file walker still
// applies those `.gitignore` files when deciding which files to scan.
if !inside_git_repo && cwd.as_ref().is_some_and(|cwd| cwd.starts_with(dir)) {
break;
}
}
// Match git's semantics: a directory is ignored when the directory itself or any
// of its parent directories is excluded, and it is not possible to re-include a
// directory once a parent directory is excluded — git never descends into an
// excluded directory, so whitelist rules inside of it are unreachable.
//
// So walk the path from the top down (`chain` is ordered bottom-up: `base` at
// index 0, the boundary last) and decide for every directory along the way
// whether it is excluded. The first excluded directory settles it. For a single
// directory, only `.gitignore` files in its parent directories can match it (its
// own `.gitignore` only matches paths _inside_ of it), and the deepest
// `.gitignore` with a definitive answer wins, so a directory that is re-included
// by a deeper `!the-directory` pattern is not ignored, even when an ancestor
// `.gitignore` ignores it.
'prefixes: for i in (0..chain.len().saturating_sub(1)).rev() {
let prefix = chain[i];
for dir in &chain[i + 1..] {
let gitignore = gitignores.entry(dir.to_path_buf()).or_insert_with(|| {
let path = dir.join(".gitignore");
// `Gitignore::new` roots the matcher at the directory containing the
// file, so patterns match relative to it.
path.is_file().then(|| Gitignore::new(&path).0)
});
let Some(gitignore) = gitignore else {
continue;
};
match gitignore.matched(prefix, true) {
ignore::Match::Ignore(_) => {
source = SourceEntry::External { base: base.into() };
break 'prefixes;
}
// Re-included; this directory is reachable, move on to the next one.
ignore::Match::Whitelist(_) => continue 'prefixes,
ignore::Match::None => {}
}
}
}
}
source
})
.collect::<Vec<SourceEntry>>();
expand_restricted_patterns(sources)
sources
}
/// Convert a public source entry to a source entry
@ -858,8 +678,15 @@ impl From<PublicSourceEntry> for SourceEntry {
};
}
let auto =
value.pattern == "/**/*" || PathBuf::from(&value.base).join(&value.pattern).is_dir();
// After a successful `optimize()` any trailing concrete directory has already been
// hoisted into the base, so a folder source always has the `/**/*` pattern. The
// `is_dir` check only matters when `optimize()` could not canonicalize the base and
// left the entry untouched. Note that the pinned leading `/` has to be stripped, since
// joining an absolute-looking path onto the base would discard the base entirely.
let auto = value.pattern == "/**/*"
|| PathBuf::from(&value.base)
.join(value.pattern.trim_start_matches('/'))
.is_dir();
if !auto {
return SourceEntry::Pattern {
@ -868,6 +695,8 @@ impl From<PublicSourceEntry> for SourceEntry {
};
}
// A directory inside e.g. `node_modules` is ignored by default, so listing it
// explicitly makes it an external source.
let inside_ignored_content_dir = IGNORED_CONTENT_DIRS.iter().any(|dir| {
value.base.contains(&format!(
"{}{}{}",
@ -889,50 +718,3 @@ impl From<PublicSourceEntry> for SourceEntry {
}
}
}
impl From<GlobEntry> for SourceEntry {
fn from(value: GlobEntry) -> Self {
SourceEntry::Pattern {
base: PathBuf::from(value.base),
pattern: value.pattern,
}
}
}
impl From<SourceEntry> for GlobEntry {
fn from(value: SourceEntry) -> Self {
match value {
SourceEntry::Auto { base } | SourceEntry::External { base } => GlobEntry {
base: base.to_string_lossy().into(),
pattern: "**/*".into(),
},
SourceEntry::Pattern { base, pattern } => GlobEntry {
base: base.to_string_lossy().into(),
pattern: pattern.clone(),
},
SourceEntry::Ignored { base, pattern } => GlobEntry {
base: base.to_string_lossy().into(),
pattern: pattern.clone(),
},
}
}
}
impl From<&SourceEntry> for GlobEntry {
fn from(value: &SourceEntry) -> Self {
match value {
SourceEntry::Auto { base } | SourceEntry::External { base } => GlobEntry {
base: base.to_string_lossy().into(),
pattern: "**/*".into(),
},
SourceEntry::Pattern { base, pattern } => GlobEntry {
base: base.to_string_lossy().into(),
pattern: pattern.clone(),
},
SourceEntry::Ignored { base, pattern } => GlobEntry {
base: base.to_string_lossy().into(),
pattern: pattern.clone(),
},
}
}
}

File diff suppressed because it is too large Load diff