Skip to content

Scanning and Filtering

Scanning walks a directory into a tree of Directory nodes; filtering decides which entries are left out along the way.

Directory

recursivist._models.Directory dataclass

A single directory within a scanned directory structure.

A scan yields a tree of these nodes: the root directory, with every subdirectory nested under its name in subdirectories. Subdirectory names are the only keys of that mapping, so a directory may be called anything the filesystem allows.

Attributes:

Name Type Description
files list[FileEntry]

The directory's own files.

subdirectories dict[str, Directory]

The directory's subdirectories, keyed by name.

loc int | None

Total lines of code in the directory and everything below it, or None when lines of code were not counted.

size int | None

Total size in bytes of the directory and everything below it, or None when sizes were not measured.

mtime float | None

Latest modification time (seconds since epoch) in the directory and everything below it, or None when modification times were not recorded.

max_depth_reached bool

Whether traversal stopped at the depth limit, leaving the directory's contents unread.

hidden_contents bool

Whether a directory cut short by the depth limit is not empty, so renderers can tell it apart from one that holds nothing.

symlink_loop bool

Whether the directory was not descended into because it resolves to one of its own ancestors, i.e. a symlink (or other) cycle back up the tree.

git_markers dict[str, str]

{filename: status_char} Git status markers for the directory's files. Empty when Git status was not collected or nothing here has a status.

Source code in recursivist/_models.py
@dataclass(slots=True)
class Directory:
    """A single directory within a scanned directory structure.

    A scan yields a tree of these nodes: the root directory, with every subdirectory
    nested under its name in ``subdirectories``. Subdirectory names are the only keys of
    that mapping, so a directory may be called anything the filesystem allows.

    Attributes:
        files: The directory's own files.
        subdirectories: The directory's subdirectories, keyed by name.
        loc: Total lines of code in the directory and everything below it, or ``None``
            when lines of code were not counted.
        size: Total size in bytes of the directory and everything below it, or ``None``
            when sizes were not measured.
        mtime: Latest modification time (seconds since epoch) in the directory and
            everything below it, or ``None`` when modification times were not recorded.
        max_depth_reached: Whether traversal stopped at the depth limit, leaving the
            directory's contents unread.
        hidden_contents: Whether a directory cut short by the depth limit is not empty,
            so renderers can tell it apart from one that holds nothing.
        symlink_loop: Whether the directory was not descended into because it resolves
            to one of its own ancestors, i.e. a symlink (or other) cycle back up the
            tree.
        git_markers: ``{filename: status_char}`` Git status markers for the directory's
            files. Empty when Git status was not collected or nothing here has a status.
    """

    files: list[FileEntry] = field(default_factory=list)
    subdirectories: dict[str, Directory] = field(default_factory=dict)
    loc: int | None = None
    size: int | None = None
    mtime: float | None = None
    max_depth_reached: bool = False
    hidden_contents: bool = False
    symlink_loop: bool = False
    git_markers: dict[str, str] = field(default_factory=dict)

FileEntry

recursivist._models.FileEntry

Bases: NamedTuple

A single file within a scanned directory structure.

Attributes:

Name Type Description
name str

Bare filename (e.g. "main.py"). Used for icon lookup, extension detection and Git-status lookup.

path str

The string to display for this file — an absolute, forward-slash path when full-path display is enabled, otherwise just name.

loc int

Lines of code. Populated only when LOC counting is enabled during scanning; 0 otherwise.

size int

File size in bytes. Populated only when size tracking is enabled; 0 otherwise.

mtime float

Modification time (seconds since epoch). Populated only when mtime tracking is enabled; 0.0 otherwise.

Source code in recursivist/_models.py
class FileEntry(NamedTuple):
    """A single file within a scanned directory structure.

    Attributes:
        name: Bare filename (e.g. ``"main.py"``). Used for icon lookup, extension
            detection and Git-status lookup.
        path: The string to display for this file — an absolute, forward-slash path when
            full-path display is enabled, otherwise just ``name``.
        loc: Lines of code. Populated only when LOC counting is enabled during scanning;
            ``0`` otherwise.
        size: File size in bytes. Populated only when size tracking is enabled; ``0``
            otherwise.
        mtime: Modification time (seconds since epoch). Populated only when mtime
            tracking is enabled; ``0.0`` otherwise.
    """

    name: str
    path: str
    loc: int = 0
    size: int = 0
    mtime: float = 0.0

Scanner

recursivist.scanner

Directory traversal.

Recursively walks a directory, applies the exclusion rules from recursivist.filtering, collects optional per-file metrics from recursivist.metrics, and returns the tree of Directory nodes consumed by the renderers and exporters.

has_contents

has_contents(directory: Directory) -> bool

Return whether a scanned directory holds anything to display.

A directory counts as non-empty when it has files, has subdirectories, links back to one of its ancestors, or was cut short by the depth limit with contents left unexplored. Renderers use this to pick between the open and closed folder icons.

Parameters:

Name Type Description Default
directory Directory

A directory from a structure produced by get_directory_structure.

required

Returns:

Type Description
bool

True if the directory has visible contents.

Source code in recursivist/scanner.py
def has_contents(directory: Directory) -> bool:
    """Return whether a scanned directory holds anything to display.

    A directory counts as non-empty when it has files, has subdirectories, links back to
    one of its ancestors, or was cut short by the depth limit with contents left
    unexplored. Renderers use this to pick between the open and closed folder icons.

    Args:
        directory: A directory from a structure produced by
            [`get_directory_structure`][recursivist.scanner.get_directory_structure].

    Returns:
        ``True`` if the directory has visible contents.
    """
    if directory.max_depth_reached:
        return directory.hidden_contents
    if directory.symlink_loop:
        return True
    return bool(directory.files or directory.subdirectories)

get_directory_structure

get_directory_structure(root_dir: str, exclude_dirs: Sequence[str] | None = None, ignore_file: str | None = None, exclude_extensions: set[str] | None = None, parent_ignore_patterns: Sequence[tuple[str, tuple[str, ...]]] | None = None, exclude_patterns: Sequence[str | Pattern[str]] | None = None, include_patterns: Sequence[str | Pattern[str]] | None = None, max_depth: int = 0, current_depth: int = 0, current_path: str = '', show_full_path: bool = False, sort_by_loc: bool = False, sort_by_size: bool = False, sort_by_mtime: bool = False, show_git_status: bool = False, git_status_map: dict[str, str] | None = None, ancestor_ids: frozenset[tuple[int, int]] | None = None, pattern_tracker: PatternMatchTracker | None = None) -> tuple[Directory, set[str]]

Build the tree of directory nodes representing a directory structure.

Recursively traverses root_dir, applying the exclusion rules and optionally collecting per-file metrics, and returns the root Directory consumed by the renderers and exporters. Each subdirectory is a nested Directory stored under its name in the parent's subdirectories; iterate them in display order with iter_subdirectories.

Fields set on each directory of the returned structure:

  • files: a FileEntry for each of the directory's files.
  • subdirectories: the nested directories, keyed by name.
  • loc: total lines of code (when sort_by_loc is set).
  • size: total size in bytes (when sort_by_size is set).
  • mtime: latest modification time (when sort_by_mtime is set).
  • max_depth_reached: True when traversal stopped at max_depth.
  • hidden_contents: True alongside max_depth_reached when the untraversed directory is not empty, so renderers can still tell it apart from one that holds nothing.
  • symlink_loop: True when a directory was not recursed into because it resolves to one of its own ancestors, i.e. a symlink (or other) cycle back up the tree.
  • git_markers: {filename: status_char} (when show_git_status is set).

A metric that was not requested is left as None, as are all three on a directory that was not traversed (one cut short by the depth limit, or a symlink loop).

When show_git_status is set, files Git reports as deleted are listed even though they are no longer on disk, together with any directory that disappeared along with them. They are subject to the same exclusion rules as every other entry, and a deleted directory is listed only if at least one of its deleted files survives them.

Parameters:

Name Type Description Default
root_dir str

Directory to scan.

required
exclude_dirs Sequence[str] | None

Directory names to skip entirely.

None
ignore_file str | None

Name of an ignore file to honor within each directory (e.g. .gitignore).

None
exclude_extensions set[str] | None

Lowercase, dot-prefixed extensions to exclude.

None
parent_ignore_patterns Sequence[tuple[str, tuple[str, ...]]] | None

Ignore files inherited from parent directories as a shallowest-first stack of (base_dir_relative_to_root, patterns) pairs. Each ignore file keeps its own anchoring so its patterns stay scoped to its subtree, matching Git. Set internally across the recursion.

None
exclude_patterns Sequence[str | Pattern[str]] | None

Glob or compiled-regex patterns to exclude.

None
include_patterns Sequence[str | Pattern[str]] | None

Glob or compiled-regex patterns to include. When given, only files whose names match one are kept, and a match overrides ignore-file rules for that file. They do not override exclude_dirs, exclude_extensions, or exclude_patterns.

None
max_depth int

Maximum depth to traverse, or 0 for unlimited.

0
current_depth int

Current recursion depth. Set internally.

0
current_path str

Path of the current directory relative to the scan root. Set internally.

''
show_full_path bool

Whether to store absolute paths instead of bare filenames.

False
sort_by_loc bool

Whether to count and total lines of code.

False
sort_by_size bool

Whether to measure and total file sizes.

False
sort_by_mtime bool

Whether to record file modification times.

False
show_git_status bool

Whether to annotate files with Git status markers.

False
git_status_map dict[str, str] | None

Pre-computed {rel_path: status_char} mapping, as returned by recursivist.git_status.get_git_status.

None
ancestor_ids frozenset[tuple[int, int]] | None

(st_dev, st_ino) identities of the directories on the path from the scan root to (and including) root_dir, used to detect symlink cycles. Set internally across the recursion.

None
pattern_tracker PatternMatchTracker | None

Records which of the exclude/include filters matched a scanned entry. When omitted on the top-level call, one is created and each filter that matched nothing is logged as a warning once the scan finishes. Pass one explicitly to aggregate several scans (as a comparison does) and report it yourself.

None

Returns:

Type Description
Directory

A (structure, extensions) tuple, where structure is the root directory

set[str]

node and extensions is the set of lowercase file extensions encountered.

Source code in recursivist/scanner.py
def get_directory_structure(
    root_dir: str,
    exclude_dirs: Sequence[str] | None = None,
    ignore_file: str | None = None,
    exclude_extensions: set[str] | None = None,
    parent_ignore_patterns: Sequence[tuple[str, tuple[str, ...]]] | None = None,
    exclude_patterns: Sequence[str | Pattern[str]] | None = None,
    include_patterns: Sequence[str | Pattern[str]] | None = None,
    max_depth: int = 0,
    current_depth: int = 0,
    current_path: str = "",
    show_full_path: bool = False,
    sort_by_loc: bool = False,
    sort_by_size: bool = False,
    sort_by_mtime: bool = False,
    show_git_status: bool = False,
    git_status_map: dict[str, str] | None = None,
    ancestor_ids: frozenset[tuple[int, int]] | None = None,
    pattern_tracker: PatternMatchTracker | None = None,
) -> tuple[Directory, set[str]]:
    """Build the tree of directory nodes representing a directory structure.

    Recursively traverses *root_dir*, applying the exclusion rules and optionally
    collecting per-file metrics, and returns the root
    [`Directory`][recursivist._models.Directory] consumed by the renderers and
    exporters. Each subdirectory is a nested
    [`Directory`][recursivist._models.Directory] stored under its name in the parent's
    ``subdirectories``; iterate them in display order with
    [`iter_subdirectories`][recursivist.scanner.iter_subdirectories].

    Fields set on each directory of the returned structure:

    - ``files``: a [`FileEntry`][recursivist._models.FileEntry] for each of the
      directory's files.
    - ``subdirectories``: the nested directories, keyed by name.
    - ``loc``: total lines of code (when *sort_by_loc* is set).
    - ``size``: total size in bytes (when *sort_by_size* is set).
    - ``mtime``: latest modification time (when *sort_by_mtime* is set).
    - ``max_depth_reached``: ``True`` when traversal stopped at *max_depth*.
    - ``hidden_contents``: ``True`` alongside ``max_depth_reached`` when the untraversed
      directory is not empty, so renderers can still tell it apart from one that holds
      nothing.
    - ``symlink_loop``: ``True`` when a directory was not recursed into because it
      resolves to one of its own ancestors, i.e. a symlink (or other) cycle back up the
      tree.
    - ``git_markers``: ``{filename: status_char}`` (when *show_git_status* is set).

    A metric that was not requested is left as ``None``, as are all three on a directory
    that was not traversed (one cut short by the depth limit, or a symlink loop).

    When *show_git_status* is set, files Git reports as deleted are listed even though
    they are no longer on disk, together with any directory that disappeared along with
    them. They are subject to the same exclusion rules as every other entry, and a
    deleted directory is listed only if at least one of its deleted files survives them.

    Args:
        root_dir: Directory to scan.
        exclude_dirs: Directory names to skip entirely.
        ignore_file: Name of an ignore file to honor within each directory (e.g.
            ``.gitignore``).
        exclude_extensions: Lowercase, dot-prefixed extensions to exclude.
        parent_ignore_patterns: Ignore files inherited from parent directories as a
            shallowest-first stack of ``(base_dir_relative_to_root, patterns)`` pairs.
            Each ignore file keeps its own anchoring so its patterns stay scoped to its
            subtree, matching Git. Set internally across the recursion.
        exclude_patterns: Glob or compiled-regex patterns to exclude.
        include_patterns: Glob or compiled-regex patterns to include. When given, only
            files whose names match one are kept, and a match overrides ignore-file
            rules for that file. They do not override *exclude_dirs*,
            *exclude_extensions*, or *exclude_patterns*.
        max_depth: Maximum depth to traverse, or ``0`` for unlimited.
        current_depth: Current recursion depth. Set internally.
        current_path: Path of the current directory relative to the scan root. Set
            internally.
        show_full_path: Whether to store absolute paths instead of bare filenames.
        sort_by_loc: Whether to count and total lines of code.
        sort_by_size: Whether to measure and total file sizes.
        sort_by_mtime: Whether to record file modification times.
        show_git_status: Whether to annotate files with Git status markers.
        git_status_map: Pre-computed ``{rel_path: status_char}`` mapping, as returned by
            [`recursivist.git_status.get_git_status`][recursivist.git_status.get_git_status].
        ancestor_ids: ``(st_dev, st_ino)`` identities of the directories on the path
            from the scan root to (and including) *root_dir*, used to detect symlink
            cycles. Set internally across the recursion.
        pattern_tracker: Records which of the exclude/include filters matched a
            scanned entry. When omitted on the top-level call, one is created and each
            filter that matched nothing is logged as a warning once the scan finishes.
            Pass one explicitly to aggregate several scans (as a comparison does) and
            report it yourself.

    Returns:
        A ``(structure, extensions)`` tuple, where *structure* is the root directory
        node and *extensions* is the set of lowercase file extensions encountered.
    """
    if exclude_dirs is None:
        exclude_dirs = []
    if exclude_extensions is None:
        exclude_extensions = set()
    if exclude_patterns is None:
        exclude_patterns = []
    if include_patterns is None:
        include_patterns = []
    if ancestor_ids is None:
        ancestor_ids = frozenset()
    owns_tracker = pattern_tracker is None and current_depth == 0
    if pattern_tracker is None:
        pattern_tracker = PatternMatchTracker(
            exclude_dirs, exclude_extensions, exclude_patterns, include_patterns
        )
    git_markers_by_dir = (
        _group_git_status(git_status_map)
        if show_git_status and git_status_map is not None
        else None
    )
    structure, extensions_set = _scan_level(
        root_dir,
        exclude_dirs,
        ignore_file,
        exclude_extensions,
        parent_ignore_patterns,
        exclude_patterns,
        include_patterns,
        max_depth,
        current_depth,
        current_path,
        show_full_path,
        sort_by_loc,
        sort_by_size,
        sort_by_mtime,
        git_markers_by_dir,
        _DeletedEntries(git_markers_by_dir) if git_markers_by_dir is not None else None,
        ancestor_ids,
        pattern_tracker,
    )
    if owns_tracker:
        pattern_tracker.report()
    return structure, extensions_set

iter_subdirectories

iter_subdirectories(directory: Directory) -> Iterator[tuple[str, Directory]]

Yield (name, subdirectory) for each subdirectory of directory.

Entries are yielded in case-sensitive name order, which is the order every renderer and exporter lists subdirectories in.

Parameters:

Name Type Description Default
directory Directory

A directory from a structure produced by get_directory_structure.

required

Yields:

Type Description
tuple[str, Directory]

(subdirectory_name, subdirectory) pairs.

Source code in recursivist/scanner.py
def iter_subdirectories(directory: Directory) -> Iterator[tuple[str, Directory]]:
    """Yield ``(name, subdirectory)`` for each subdirectory of *directory*.

    Entries are yielded in case-sensitive name order, which is the order every renderer
    and exporter lists subdirectories in.

    Args:
        directory: A directory from a structure produced by
            [`get_directory_structure`][recursivist.scanner.get_directory_structure].

    Yields:
        ``(subdirectory_name, subdirectory)`` pairs.
    """
    yield from sorted(directory.subdirectories.items(), key=lambda entry: entry[0])

collect_extensions

collect_extensions(directory: Directory) -> set[str]

Return the lowercase file extensions of every file in directory.

Walks the whole structure and gathers the same set that get_directory_structure returns alongside it, for callers that hold only the structure.

Parameters:

Name Type Description Default
directory Directory

The root of a structure produced by get_directory_structure.

required

Returns:

Type Description
set[str]

The set of extensions, each lowercase with its leading dot (e.g. ".py").

set[str]

Files without an extension contribute nothing.

Source code in recursivist/scanner.py
def collect_extensions(directory: Directory) -> set[str]:
    """Return the lowercase file extensions of every file in *directory*.

    Walks the whole structure and gathers the same set that
    [`get_directory_structure`][recursivist.scanner.get_directory_structure] returns
    alongside it, for callers that hold only the structure.

    Args:
        directory: The root of a structure produced by
            [`get_directory_structure`][recursivist.scanner.get_directory_structure].

    Returns:
        The set of extensions, each lowercase with its leading dot (e.g. ``".py"``).
        Files without an extension contribute nothing.
    """
    extensions: set[str] = set()
    for entry in directory.files:
        _, ext = os.path.splitext(entry.name)
        if ext:
            extensions.add(ext.lower())
    for subdirectory in directory.subdirectories.values():
        extensions.update(collect_extensions(subdirectory))
    return extensions

Filtering

recursivist.filtering

File and directory filtering: ignore files, glob/regex patterns, and gitignore-style exclusion rules.

Provides the predicate should_exclude used by the scanner. Git-style ignore matching is delegated to pathspec (its gitignore matcher), which implements the full gitignore specification: anchoring, ** wildcards, directory-only (trailing /) patterns, ! negation with last-match-wins, character classes, backslash escapes, and trailing-whitespace handling. The glob and regex matching used by --exclude-pattern/--include-pattern is unrelated and uses only the standard library.

Like Git, each ignore file is evaluated relative to the directory that contains it rather than relative to the scan root. The active ignore files are kept as a stack (shallowest first); a path is tested against every level with its own anchoring, and a deeper file's verdict overrides a shallower one, so an anchored pattern such as /build in a nested .gitignore matches only within that subdirectory and does not leak up to the scan root or down past the anchor.

InvalidPatternError

Bases: ValueError

Raised when a --regex pattern is not a valid regular expression.

Source code in recursivist/filtering.py
class InvalidPatternError(ValueError):
    """Raised when a ``--regex`` pattern is not a valid regular expression."""

PatternMatchTracker

Record which user-supplied filters matched at least one scanned entry.

The scanner reports every directory entry it lists to observe before applying the exclusion rules, so a filter counts as matched even when an earlier rule already removed the entry. Once a filter has matched it is dropped from the pending set, and once nothing is pending observation is a no-op, so the bookkeeping costs nothing on a scan where every filter is used.

Matching mirrors the scanner exactly: --exclude names are compared with the entry name, --exclude-ext applies to files only, --exclude-pattern applies to files and directories, and --include-pattern applies to files only.

Entries the scanner never lists — the contents of excluded or ignored directories, and anything below --depth — are not observed, so a filter reported as unmatched matched nothing among the scanned entries.

Parameters:

Name Type Description Default
exclude_dirs Sequence[str] | None

Names given to --exclude.

None
exclude_extensions Iterable[str] | None

Normalized extensions given to --exclude-ext.

None
exclude_patterns Sequence[str | Pattern[str]] | None

Patterns given to --exclude-pattern.

None
include_patterns Sequence[str | Pattern[str]] | None

Patterns given to --include-pattern.

None
Source code in recursivist/filtering.py
class PatternMatchTracker:
    """Record which user-supplied filters matched at least one scanned entry.

    The scanner reports every directory entry it lists to
    [`observe`][recursivist.filtering.PatternMatchTracker.observe] *before* applying the
    exclusion rules, so a filter counts as matched even when an earlier rule already
    removed the entry. Once a filter has matched it is dropped from the pending set, and
    once nothing is pending observation is a no-op, so the bookkeeping costs nothing on
    a scan where every filter is used.

    Matching mirrors the scanner exactly: ``--exclude`` names are compared with the
    entry name, ``--exclude-ext`` applies to files only, ``--exclude-pattern`` applies
    to files and directories, and ``--include-pattern`` applies to files only.

    Entries the scanner never lists — the contents of excluded or ignored directories,
    and anything below ``--depth`` — are not observed, so a filter reported as
    unmatched matched nothing *among the scanned entries*.

    Args:
        exclude_dirs: Names given to ``--exclude``.
        exclude_extensions: Normalized extensions given to ``--exclude-ext``.
        exclude_patterns: Patterns given to ``--exclude-pattern``.
        include_patterns: Patterns given to ``--include-pattern``.
    """

    def __init__(
        self,
        exclude_dirs: Sequence[str] | None = None,
        exclude_extensions: Iterable[str] | None = None,
        exclude_patterns: Sequence[str | Pattern[str]] | None = None,
        include_patterns: Sequence[str | Pattern[str]] | None = None,
    ) -> None:
        self._dirs: dict[str, None] = dict.fromkeys(exclude_dirs or ())
        self._exts: dict[str, None] = dict.fromkeys(exclude_extensions or ())
        self._exclude: list[str | Pattern[str]] = list(
            dict.fromkeys(exclude_patterns or ())
        )
        self._include: list[str | Pattern[str]] = list(
            dict.fromkeys(include_patterns or ())
        )
        self.depth_limited = False

    @property
    def pending(self) -> bool:
        """Whether any filter has not matched an entry yet."""
        return bool(self._dirs or self._exts or self._exclude or self._include)

    def observe(self, path: str, is_dir: bool) -> None:
        """Mark every pending filter that matches the entry at *path* as used.

        Args:
            path: Filesystem path of a directory entry the scanner listed.
            is_dir: Whether *path* is a directory.
        """
        if not self.pending:
            return
        name = os.path.basename(path)
        self._dirs.pop(name, None)
        if not is_dir and self._exts:
            self._exts.pop(os.path.splitext(name)[1].lower(), None)
        if self._exclude:
            self._exclude = [p for p in self._exclude if not _pattern_matches(p, name)]
        if not is_dir and self._include:
            self._include = [p for p in self._include if not _pattern_matches(p, name)]

    def unmatched(self) -> list[tuple[str, str]]:
        """Return ``(flag, value)`` pairs for every filter that matched nothing."""
        return [
            *(("--exclude", d) for d in self._dirs),
            *(("--exclude-ext", e) for e in self._exts),
            *(("--exclude-pattern", _describe_pattern(p)) for p in self._exclude),
            *(("--include-pattern", _describe_pattern(p)) for p in self._include),
        ]

    def report(self, where: str = "") -> None:
        """Log a warning for each filter that matched nothing.

        Args:
            where: Optional phrase naming what was scanned (e.g. ``"in either
                directory"``), appended to each message.
        """
        suffix = f" {where}" if where else ""
        if self.depth_limited:
            suffix += " within the scanned depth"
        for flag, value in self.unmatched():
            logger.warning(
                "No files or directories matched %s '%s'%s", flag, value, suffix
            )

pending property

pending: bool

Whether any filter has not matched an entry yet.

observe

observe(path: str, is_dir: bool) -> None

Mark every pending filter that matches the entry at path as used.

Parameters:

Name Type Description Default
path str

Filesystem path of a directory entry the scanner listed.

required
is_dir bool

Whether path is a directory.

required
Source code in recursivist/filtering.py
def observe(self, path: str, is_dir: bool) -> None:
    """Mark every pending filter that matches the entry at *path* as used.

    Args:
        path: Filesystem path of a directory entry the scanner listed.
        is_dir: Whether *path* is a directory.
    """
    if not self.pending:
        return
    name = os.path.basename(path)
    self._dirs.pop(name, None)
    if not is_dir and self._exts:
        self._exts.pop(os.path.splitext(name)[1].lower(), None)
    if self._exclude:
        self._exclude = [p for p in self._exclude if not _pattern_matches(p, name)]
    if not is_dir and self._include:
        self._include = [p for p in self._include if not _pattern_matches(p, name)]

unmatched

unmatched() -> list[tuple[str, str]]

Return (flag, value) pairs for every filter that matched nothing.

Source code in recursivist/filtering.py
def unmatched(self) -> list[tuple[str, str]]:
    """Return ``(flag, value)`` pairs for every filter that matched nothing."""
    return [
        *(("--exclude", d) for d in self._dirs),
        *(("--exclude-ext", e) for e in self._exts),
        *(("--exclude-pattern", _describe_pattern(p)) for p in self._exclude),
        *(("--include-pattern", _describe_pattern(p)) for p in self._include),
    ]

report

report(where: str = '') -> None

Log a warning for each filter that matched nothing.

Parameters:

Name Type Description Default
where str

Optional phrase naming what was scanned (e.g. "in either directory"), appended to each message.

''
Source code in recursivist/filtering.py
def report(self, where: str = "") -> None:
    """Log a warning for each filter that matched nothing.

    Args:
        where: Optional phrase naming what was scanned (e.g. ``"in either
            directory"``), appended to each message.
    """
    suffix = f" {where}" if where else ""
    if self.depth_limited:
        suffix += " within the scanned depth"
    for flag, value in self.unmatched():
        logger.warning(
            "No files or directories matched %s '%s'%s", flag, value, suffix
        )

parse_ignore_file

parse_ignore_file(ignore_file_path: str) -> list[str]

Read an ignore file and return its lines as gitignore patterns.

Lines are returned verbatim with only their terminators removed, preserving order and every character that is significant to the gitignore grammar (comments, blank lines, backslash escapes, and escaped trailing whitespace). Interpretation is left entirely to the gitignore matcher, so callers must not strip or filter the returned lines.

A leading UTF-8 byte order mark is dropped, as Git does, so that it does not become part of the first pattern and silently stop that pattern from matching.

Parameters:

Name Type Description Default
ignore_file_path str

Path to the ignore file (e.g. .gitignore).

required

Returns:

Type Description
list[str]

The list of pattern lines, or an empty list when the path does not exist, is

list[str]

not a regular file (a directory, or a named pipe that would block when read),

list[str]

or cannot be read.

Source code in recursivist/filtering.py
def parse_ignore_file(ignore_file_path: str) -> list[str]:
    """Read an ignore file and return its lines as gitignore patterns.

    Lines are returned verbatim with only their terminators removed, preserving order
    and every character that is significant to the gitignore grammar (comments, blank
    lines, backslash escapes, and escaped trailing whitespace). Interpretation is left
    entirely to the gitignore matcher, so callers must not strip or filter the returned
    lines.

    A leading UTF-8 byte order mark is dropped, as Git does, so that it does not become
    part of the first pattern and silently stop that pattern from matching.

    Args:
        ignore_file_path: Path to the ignore file (e.g. ``.gitignore``).

    Returns:
        The list of pattern lines, or an empty list when the path does not exist, is
        not a regular file (a directory, or a named pipe that would block when read),
        or cannot be read.
    """
    if not os.path.isfile(ignore_file_path):
        return []
    try:
        with open(ignore_file_path, encoding="utf-8-sig", errors="replace") as f:
            return f.read().splitlines()
    except OSError as e:
        logger.warning(
            "Skipping unreadable ignore file %s: %s", ignore_file_path, e.strerror or e
        )
        return []

normalize_extensions

normalize_extensions(extensions: Iterable[str]) -> set[str]

Return extensions in the lowercase, dot-prefixed form the scanner matches on.

An extension may be written with or without its leading dot and in any case, so "pyc", ".pyc" and ".PYC" all become ".pyc". The result is the form expected wherever an exclude_extensions argument is described as normalized. Duplicates collapse, and normalizing an already-normalized set returns an equal set.

Source code in recursivist/filtering.py
def normalize_extensions(extensions: Iterable[str]) -> set[str]:
    """Return *extensions* in the lowercase, dot-prefixed form the scanner matches on.

    An extension may be written with or without its leading dot and in any case, so
    ``"pyc"``, ``".pyc"`` and ``".PYC"`` all become ``".pyc"``. The result is the form
    expected wherever an ``exclude_extensions`` argument is described as normalized.
    Duplicates collapse, and normalizing an already-normalized set returns an equal
    set.
    """
    return {
        ext.lower() if ext.startswith(".") else f".{ext.lower()}" for ext in extensions
    }

compile_regex_patterns

compile_regex_patterns(patterns: Sequence[str], is_regex: bool = False) -> list[str | Pattern[str]]

Compile patterns to regex objects when regex matching is requested.

When is_regex is False the patterns are returned unchanged for glob matching. When True each pattern is compiled to a re.Pattern, and a pattern that fails to compile is an error.

Parameters:

Name Type Description Default
patterns Sequence[str]

Patterns to process.

required
is_regex bool

Whether to treat the patterns as regular expressions (True) or glob patterns (False).

False

Returns:

Type Description
list[str | Pattern[str]]

A list whose items are plain strings for glob patterns or compiled re.Pattern

list[str | Pattern[str]]

objects for regexes.

Raises:

Type Description
InvalidPatternError

If is_regex is True and any pattern is not a valid regular expression. The message names every invalid pattern.

Source code in recursivist/filtering.py
def compile_regex_patterns(
    patterns: Sequence[str], is_regex: bool = False
) -> list[str | Pattern[str]]:
    """Compile patterns to regex objects when regex matching is requested.

    When *is_regex* is ``False`` the patterns are returned unchanged for glob matching.
    When ``True`` each pattern is compiled to a `re.Pattern`, and a pattern that fails
    to compile is an error.

    Args:
        patterns: Patterns to process.
        is_regex: Whether to treat the patterns as regular expressions (``True``) or
            glob patterns (``False``).

    Returns:
        A list whose items are plain strings for glob patterns or compiled `re.Pattern`
        objects for regexes.

    Raises:
        InvalidPatternError: If *is_regex* is ``True`` and any pattern is not a valid
            regular expression. The message names every invalid pattern.
    """
    if not is_regex:
        return cast(list[str | Pattern[str]], patterns)
    compiled_patterns: list[str | Pattern[str]] = []
    errors: list[str] = []
    for pattern in patterns:
        try:
            compiled_patterns.append(re.compile(pattern))
        except re.error as e:
            errors.append(f"'{pattern}' ({e})")
    if errors:
        noun = "pattern" if len(errors) == 1 else "patterns"
        raise InvalidPatternError(f"Invalid regex {noun}: {', '.join(errors)}")
    return compiled_patterns

should_exclude

should_exclude(path: str, ignore_context: dict[str, Any], exclude_extensions: set[str] | None = None, exclude_patterns: Sequence[str | Pattern[str]] | None = None, include_patterns: Sequence[str | Pattern[str]] | None = None, *, is_dir: bool | None = None) -> bool

Decide whether a path should be excluded from the scan.

The filtering rules are applied in priority order:

  1. If include_patterns are given and none match a file, exclude it.
  2. If any exclude_patterns match, exclude the path (this overrides include patterns).
  3. If a non-directory's extension is in exclude_extensions, exclude it.
  4. If an include pattern matched a file, include it (this overrides the gitignore-style patterns below). Directories are never tested against include patterns, so they always fall through to the ignore rules.
  5. Otherwise apply the gitignore-style rules from ignore_context via pathspec, honoring anchoring, ** wildcards, directory-only (trailing /) patterns, ! negation, character classes and escapes per the gitignore specification. Each ignore file in the stack is matched relative to its own directory and deeper files override shallower ones, so a nested file's anchored patterns stay scoped to its subtree.

Parameters:

Name Type Description Default
path str

Filesystem path to test.

required
ignore_context dict[str, Any]

Mapping describing the active ignore rules. Recognized keys are "pattern_stack" (a shallowest-first sequence of (base_dir_relative_to_root, patterns) pairs, one per ignore file; a single ignore file at the scan root is [("", patterns)]) and "rel_dir" (the current directory's path relative to the scan root).

required
exclude_extensions set[str] | None

Lowercase, dot-prefixed extensions to exclude.

None
exclude_patterns Sequence[str | Pattern[str]] | None

Glob or compiled-regex patterns to exclude, matched against the entry's name.

None
include_patterns Sequence[str | Pattern[str]] | None

Glob or compiled-regex patterns to include, matched against the entry's name, which override the gitignore-style exclusions.

None
is_dir bool | None

Whether path is a directory, when the caller already knows. Left as None, it is looked up with os.path.isdir.

None

Returns:

Type Description
bool

True if the path should be excluded, False otherwise.

Source code in recursivist/filtering.py
def should_exclude(
    path: str,
    ignore_context: dict[str, Any],
    exclude_extensions: set[str] | None = None,
    exclude_patterns: Sequence[str | Pattern[str]] | None = None,
    include_patterns: Sequence[str | Pattern[str]] | None = None,
    *,
    is_dir: bool | None = None,
) -> bool:
    """Decide whether a path should be excluded from the scan.

    The filtering rules are applied in priority order:

    1. If *include_patterns* are given and none match a file, exclude it.
    2. If any *exclude_patterns* match, exclude the path (this overrides include
       patterns).
    3. If a non-directory's extension is in *exclude_extensions*, exclude it.
    4. If an include pattern matched a file, include it (this overrides the
       gitignore-style patterns below). Directories are never tested against include
       patterns, so they always fall through to the ignore rules.
    5. Otherwise apply the gitignore-style rules from *ignore_context* via `pathspec`,
       honoring anchoring, ``**`` wildcards, directory-only (trailing ``/``) patterns,
       ``!`` negation, character classes and escapes per the gitignore specification.
       Each ignore file in the stack is matched relative to its own directory and deeper
       files override shallower ones, so a nested file's anchored patterns stay scoped
       to its subtree.

    Args:
        path: Filesystem path to test.
        ignore_context: Mapping describing the active ignore rules. Recognized keys are
            ``"pattern_stack"`` (a shallowest-first sequence of
            ``(base_dir_relative_to_root, patterns)`` pairs, one per ignore file; a
            single ignore file at the scan root is ``[("", patterns)]``) and
            ``"rel_dir"`` (the current directory's path relative to the scan root).
        exclude_extensions: Lowercase, dot-prefixed extensions to exclude.
        exclude_patterns: Glob or compiled-regex patterns to exclude, matched against
            the entry's name.
        include_patterns: Glob or compiled-regex patterns to include, matched against
            the entry's name, which override the gitignore-style exclusions.
        is_dir: Whether *path* is a directory, when the caller already knows. Left as
            ``None``, it is looked up with `os.path.isdir`.

    Returns:
        ``True`` if the path should be excluded, ``False`` otherwise.
    """
    basename = os.path.basename(path)
    if is_dir is None:
        is_dir = os.path.isdir(path)
    if (
        include_patterns
        and not is_dir
        and not any(_pattern_matches(pattern, basename) for pattern in include_patterns)
    ):
        return True
    if exclude_patterns and any(
        _pattern_matches(pattern, basename) for pattern in exclude_patterns
    ):
        return True
    if (
        exclude_extensions
        and not is_dir
        and os.path.splitext(basename)[1].lower() in exclude_extensions
    ):
        return True
    if include_patterns and not is_dir:
        return False
    levels = _resolve_ignore_levels(ignore_context)
    if not levels:
        return False
    rel_dir = ignore_context.get("rel_dir", "")
    target = (f"{rel_dir}/{basename}" if rel_dir else basename).replace("\\", "/")
    target = target.lstrip("/")
    if is_dir:
        target += "/"
    return _is_ignored_by_stack(target, levels)