Fix trailing ) from interfering with extraction in Clojure keywords (#18345)

## Summary

In a form like, 

```clojure 
(if condition :bg-white :bg-black)
```

`:bg-black` will fail to extract, while `:bg-white` is extracted as
expected. This PR fixes this case, implements more comprehensive
candidate filtering, and supersedes a previous PR.

Having recently submitted a PR for handling another special case with
Clojure keywords (the presence of `:` inside of keywords), I thought it
best to invert the previous strategy: Instead of handling special cases
one by one, consume keywords according to the Clojure reader spec.
Consume nothing else, other than strings.

Because of this, this PR is a tad more invasive rather than additive,
for which I apologize. The strategy is this:
- Strings begin with a `"` and ends with an unescaped `"`. Consume
everything between these delimiters (existing case).
- Keywords begin with `:`, and end with whitespace, or one out of a
small set of specific reserved characters. Everything else is a valid
character in a keyword. Consume everything between these delimiters, and
apply the class splitting previously contained in the outer loop. My
previous special case handling of `:` inside of keywords in #18338 is
now redundant (and is removed), as this is a more general solution.
- Discard _everything else_. 

I'm hoping that a strategy that is based on Clojure's definition of
strings and keywords will pre-empt any further issues with edge cases.

Closes #18344.

## Test plan
- Added failing tests.
- `cargo test` -> failure
- Added fix
- `cargo test` -> success

---------

Co-authored-by: Jordan Pittman <jordan@cryptica.me>
This commit is contained in:
Henrik Eneroth 2025-06-30 18:23:38 +02:00 • committed by GitHub
parent 7946db05ef
commit 05b65d59b5
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
2 changed files with 112 additions and 32 deletions

View file

@ -10,6 +10,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed
- Don't consider the global important state in `@apply` ([#18404](https://github.com/tailwindlabs/tailwindcss/pull/18404))
- Fix trailing `)` from interfering with extraction in Clojure keywords ([#18345](https://github.com/tailwindlabs/tailwindcss/pull/18345))
## [4.1.11] - 2025-06-26

View file

@ -5,6 +5,23 @@ use bstr::ByteSlice;
#[derive(Debug, Default)]
pub struct Clojure;
/// This is meant to be a rough estimate of a valid ClojureScript keyword
///
/// This can be approximated by the following regex:
/// /::?[a-zA-Z0-9!#$%&*+./:<=>?_|-]+/
///
/// However, keywords are intended to be detected as utilities. Since the set
/// of valid characters in a utility (outside of arbitrary values) is smaller,
/// along with the fact that neither `[]` nor `()` are allowed in keywords we
/// can simplify this list quite a bit.
#[inline]
fn is_keyword_character(byte: u8) -> bool {
return matches!(
byte,
b'!' | b'%' | b'*' | b'+' | b'-' | b'.' | b'/' | b':' | b'_'
) | byte.is_ascii_alphanumeric();
}
impl PreProcessor for Clojure {
fn process(&self, content: &[u8]) -> Vec<u8> {
let content = content
@ -18,6 +35,7 @@ impl PreProcessor for Clojure {
match cursor.curr {
// Consume strings as-is
b'"' => {
result[cursor.pos] = b' ';
cursor.advance();
while cursor.pos < len {
@ -26,7 +44,10 @@ impl PreProcessor for Clojure {
b'\\' => cursor.advance_twice(),
// End of the string
b'"' => break,
b'"' => {
result[cursor.pos] = b' ';
break;
}
// Everything else is valid
_ => cursor.advance(),
@ -34,44 +55,71 @@ impl PreProcessor for Clojure {
}
}
// Consume comments as-is until the end of the line.
// Discard line comments until the end of the line.
// Comments start with `;;`
b';' if matches!(cursor.next, b';') => {
while cursor.pos < len && cursor.curr != b'\n' {
result[cursor.pos] = b' ';
cursor.advance();
}
}
// A `.` surrounded by digits is a decimal number, so we don't want to replace it.
//
// E.g.:
// ```
// gap-1.5
// ^
// ``
b'.' if cursor.prev.is_ascii_digit() && cursor.next.is_ascii_digit() => {
// Consume keyword until a terminating character is reached.
b':' => {
result[cursor.pos] = b' ';
cursor.advance();
// Keep the `.` as-is
while cursor.pos < len {
match cursor.curr {
// A `.` surrounded by digits is a decimal number, so we don't want to replace it.
//
// E.g.:
// ```
// gap-1.5
// ^
// ```
b'.' if cursor.prev.is_ascii_digit()
&& cursor.next.is_ascii_digit() =>
{
// Keep the `.` as-is
}
// A `.` not surrounded by digits denotes the start of a new class name in a
// dot-delimited keyword.
//
// E.g.:
// ```
// flex.gap-1.5
// ^
// ```
b'.' => {
result[cursor.pos] = b' ';
}
// End of keyword.
_ if !is_keyword_character(cursor.curr) => {
result[cursor.pos] = b' ';
break;
}
// Consume everything else.
_ => {}
};
cursor.advance();
}
}
// A `:` surrounded by letters denotes a variant. Keep as is.
//
// Aggressively discard everything else, reducing false positives and preventing
// characters surrounding keywords from producing false negatives.
// E.g.:
// ```
// lg:pr-6"
// ^
// ``
b':' if cursor.prev.is_ascii_alphanumeric() && cursor.next.is_ascii_alphanumeric() => {
// Keep the `:` as-is
}
b':' | b'.' => {
// (when condition :bg-white)
// ^
// ```
// A ')' is never a valid part of a keyword, but will nonetheless prevent 'bg-white'
// from being extracted if not discarded.
_ => {
result[cursor.pos] = b' ';
}
// Consume everything else
_ => {}
};
cursor.advance();
@ -92,19 +140,23 @@ mod tests {
(":div.flex-1.flex-2", " div flex-1 flex-2"),
(
":.flex-3.flex-4 ;defaults to div",
" flex-3 flex-4 ;defaults to div",
" flex-3 flex-4 ",
),
("{:class :flex-5.flex-6", "{ flex-5 flex-6"),
(r#"{:class "flex-7 flex-8"}"#, r#"{ "flex-7 flex-8"}"#),
("{:class :flex-5.flex-6", " flex-5 flex-6"),
(r#"{:class "flex-7 flex-8"}"#, r#" flex-7 flex-8 "#),
(
r#"{:class ["flex-9" :flex-10]}"#,
r#"{ ["flex-9" flex-10]}"#,
r#" flex-9 flex-10 "#,
),
(
r#"(dom/div {:class "flex-11 flex-12"})"#,
r#"(dom/div { "flex-11 flex-12"})"#,
r#" flex-11 flex-12 "#,
),
("(dom/div :.flex-13.flex-14", " flex-13 flex-14"),
(
r#"[:div#hello.bg-white.pr-1.5 {:class ["grid grid-cols-[auto,1fr] grid-rows-2"]}]"#,
r#" div#hello bg-white pr-1.5 grid grid-cols-[auto,1fr] grid-rows-2 "#,
),
("(dom/div :.flex-13.flex-14", "(dom/div flex-13 flex-14"),
] {
Clojure::test(input, expected);
}
@ -198,8 +250,35 @@ mod tests {
($ :div {:class [:flex :first:lg:pr-6 :first:2xl:pl-6 :group-hover/2:2xs:pt-6]} …)
:.hover:bg-white
[:div#hello.bg-white.pr-1.5]
"#;
Clojure::test_extract_contains(input, vec!["flex", "first:lg:pr-6", "first:2xl:pl-6", "group-hover/2:2xs:pt-6", "hover:bg-white"]);
Clojure::test_extract_contains(
input,
vec![
"flex",
"first:lg:pr-6",
"first:2xl:pl-6",
"group-hover/2:2xs:pt-6",
"hover:bg-white",
"bg-white",
"pr-1.5",
],
);
}
// https://github.com/tailwindlabs/tailwindcss/issues/18344
#[test]
fn test_noninterference_of_parens_on_keywords() {
let input = r#"
(get props :y-padding :py-5)
($ :div {:class [:flex.pr-1.5 (if condition :bg-white :bg-black)]})
"#;
Clojure::test_extract_contains(
input,
vec!["py-5", "flex", "pr-1.5", "bg-white", "bg-black"],
);
}
}