Skip to content

ontolith.query

ontolith.query

Query builder for traversal and retrieval.

This module provides a fluent query interface for retrieving entities and traversing the knowledge graph.

QueryBuilder

QueryBuilder(
    backend: StorageBackend,
    namespace: str,
    concept: str,
    as_of_time: datetime | None = None,
    embedder: Embedder | None = None,
)

Fluent query interface for entities.

Example

kb.query(Person).where(name="Ada Lovelace") kb.as_of("2025-01-01").query(Person).where(employer="org-123") kb.query(Person).semantic("a computer scientist").limit(5).all()

Initialize query builder.

Parameters:

Name Type Description Default
backend StorageBackend

Storage backend for retrieval

required
namespace str

Namespace to query in

required
concept str

Concept to filter by

required
as_of_time datetime | None

If set, applies bitemporal filter to all queries

None
embedder Embedder | None

Embedder used to embed .semantic() query text. Required only if .semantic() is called.

None
Source code in src/ontolith/query/builder.py
def __init__(
    self,
    backend: StorageBackend,
    namespace: str,
    concept: str,
    as_of_time: datetime | None = None,
    embedder: Embedder | None = None,
) -> None:
    """Initialize query builder.

    Args:
        backend: Storage backend for retrieval
        namespace: Namespace to query in
        concept: Concept to filter by
        as_of_time: If set, applies bitemporal filter to all queries
        embedder: Embedder used to embed .semantic() query text. Required
            only if .semantic() is called.
    """
    self._backend = backend
    self._namespace = namespace
    self._concept = concept
    self._filters: dict[tuple[str, str], Any] = {}
    self._as_of_time = as_of_time
    self._embedder = embedder
    self._semantic_text: str | None = None
    self._min_confidence: float | None = None
    self._trust_at_least: int | None = None
    self._limit: int | None = None
    self._include_flagged: bool = False
    self._include_history: bool = False

where

where(**kwargs: Any) -> QueryBuilder

Add filters to the query.

Plain keys are equality checks against the concept's own predicates — both literal properties (name="Ada Lovelace") and relations, where the value is compared against the relation's target entity id (employer="org-123"). A closed set of double-underscore ("dunder") lookup-operator suffixes is also recognized (KI-039): __contains (substring match, literal properties only) and __gt/__lt/__gte/__lte (numeric range, restricted to predicates the active schema declares Integer or Float — see _RANGE_VALUE_TYPES's docstring for why). Calling .where() more than once with the same key AND the same operator overwrites the earlier value (last call wins), matching plain equality's existing behavior; different operators on the same property compose as AND (e.g. .where(age__gte=18, age__lt=65)). Multi-hop relation traversal (a hypothetical employer__name=, ADR-0027) and any dunder suffix outside this closed operator set are still rejected outright — they never matched anything before KI-030 fixed that.

Parameters:

Name Type Description Default
**kwargs Any

Property/relation filters as keyword arguments, optionally suffixed with a recognized lookup operator

{}

Returns:

Type Description
QueryBuilder

Self for chaining

Raises:

Type Description
ValidationError

A filter key uses relation-traversal or an unrecognized dunder suffix, or a __gt/__lt/__gte/ __lte filter targets a non-numeric (or relation, or undeclared) predicate, or its value doesn't parse as a number.

Source code in src/ontolith/query/builder.py
def where(self, **kwargs: Any) -> "QueryBuilder":
    """Add filters to the query.

    Plain keys are equality checks against the concept's own predicates
    — both literal properties (`name="Ada Lovelace"`) and relations,
    where the value is compared against the relation's target entity id
    (`employer="org-123"`). A closed set of double-underscore ("dunder")
    lookup-operator suffixes is also recognized (KI-039):
    `__contains` (substring match, literal properties only) and
    `__gt`/`__lt`/`__gte`/`__lte` (numeric range, restricted to
    predicates the active schema declares `Integer` or `Float` — see
    `_RANGE_VALUE_TYPES`'s docstring for why). Calling `.where()` more
    than once with the same key AND the same operator overwrites the
    earlier value (last call wins), matching plain equality's existing
    behavior; different operators on the same property compose as AND
    (e.g. `.where(age__gte=18, age__lt=65)`). Multi-hop relation
    traversal (a hypothetical `employer__name=`, ADR-0027) and any
    dunder suffix outside this closed operator set are still rejected
    outright — they never matched anything before KI-030 fixed that.

    Args:
        **kwargs: Property/relation filters as keyword arguments,
            optionally suffixed with a recognized lookup operator

    Returns:
        Self for chaining

    Raises:
        ValidationError: A filter key uses relation-traversal or an
            unrecognized dunder suffix, or a `__gt`/`__lt`/`__gte`/
            `__lte` filter targets a non-numeric (or relation, or
            undeclared) predicate, or its value doesn't parse as a
            number.
    """
    for key, value in kwargs.items():
        prop, operator = self._parse_filter_key(key)
        if operator in _RANGE_OPERATORS:
            value = self._coerce_range_value(prop, operator, value)
        elif operator == "contains" and not isinstance(value, str):
            raise ValidationError(f"where({prop}__contains={value!r}) requires a string value")
        self._filters[(prop, operator)] = value
    return self

semantic

semantic(text: str) -> QueryBuilder

Rank results by vector similarity to text (SPEC §11.3).

Results are embedded entities (see Ontology.reindex()) ranked by ascending distance to text's embedding. Combine with .where() to intersect with symbolic filters, preserving vector rank order.

Combined with .as_of(): an entity that didn't exist yet at that point in time is excluded (KI-058), but the ranking itself is not bitemporal — the vector index holds one embedding per entity, as of whenever Ontology.reindex() was last called, with no historical versions. Results are always ranked by an entity's current embedded content, never its content as it stood at as_of_time.

Parameters:

Name Type Description Default
text str

Query text, embedded via this builder's Embedder.

required

Returns:

Type Description
QueryBuilder

Self for chaining

Source code in src/ontolith/query/builder.py
def semantic(self, text: str) -> "QueryBuilder":
    """Rank results by vector similarity to `text` (SPEC §11.3).

    Results are embedded entities (see `Ontology.reindex()`) ranked by
    ascending distance to `text`'s embedding. Combine with `.where()` to
    intersect with symbolic filters, preserving vector rank order.

    Combined with `.as_of()`: an entity that didn't exist yet at that
    point in time is excluded (KI-058), but the ranking itself is not
    bitemporal — the vector index holds one embedding per entity, as of
    whenever `Ontology.reindex()` was last called, with no historical
    versions. Results are always ranked by an entity's *current*
    embedded content, never its content as it stood at `as_of_time`.

    Args:
        text: Query text, embedded via this builder's Embedder.

    Returns:
        Self for chaining
    """
    self._semantic_text = text
    return self

min_confidence

min_confidence(threshold: float) -> QueryBuilder

Keep only entities with at least one qualifying assertion at or above threshold confidence.

Current-state (no .as_of()): active by default, also flagged when .include_flagged() is set and/or superseded/retracted when .include_history() is set (KI-093). Under .as_of(t), the qualifying set is never restricted to active in the first place — it already includes a superseded assertion valid at t regardless of status now, so .include_history() has nothing to add there. .include_flagged() still applies (excludes flagged unless set, reconstructed point-in-time from the event log, KI-097 — same as .where()'s own as_of handling). A retracted assertion is the one exception (ADR-0049, KI-095): as_of(t) additionally excludes it once its own retraction event's assertion-time has passed, and .include_history() is what opts back into seeing it.

An entity with only confidence=None assertions does not pass — None never satisfies a numeric threshold (ADR-0004). Independent of .trust_at_least(): the qualifying assertion need not be the same one for both filters. Respects .as_of() (KI-036): if this query is pinned to a point in time, the qualifying assertion must have been valid at that time, not merely currently active.

Parameters:

Name Type Description Default
threshold float

Minimum confidence, 0.0-1.0.

required

Returns:

Type Description
QueryBuilder

Self for chaining

Source code in src/ontolith/query/builder.py
def min_confidence(self, threshold: float) -> "QueryBuilder":
    """Keep only entities with at least one qualifying assertion at or
    above `threshold` confidence.

    Current-state (no `.as_of()`): `active` by default, also `flagged`
    when `.include_flagged()` is set and/or `superseded`/`retracted`
    when `.include_history()` is set (KI-093). Under `.as_of(t)`, the
    qualifying set is never restricted to `active` in the first place —
    it already includes a `superseded` assertion valid at `t` regardless
    of status now, so `.include_history()` has nothing to add *there*.
    `.include_flagged()` still applies (excludes `flagged` unless set,
    reconstructed point-in-time from the event log, KI-097 — same as
    `.where()`'s own `as_of` handling). A `retracted` assertion is the
    one exception (ADR-0049, KI-095): `as_of(t)` additionally excludes
    it once its own retraction event's assertion-time has passed, and
    `.include_history()` is what opts back into seeing it.

    An entity with only `confidence=None` assertions does not pass —
    None never satisfies a numeric threshold (ADR-0004). Independent of
    `.trust_at_least()`: the qualifying assertion need not be the same
    one for both filters. Respects `.as_of()` (KI-036): if this query
    is pinned to a point in time, the qualifying assertion must have
    been valid at that time, not merely currently active.

    Args:
        threshold: Minimum confidence, 0.0-1.0.

    Returns:
        Self for chaining
    """
    self._min_confidence = threshold
    return self

trust_at_least

trust_at_least(level: int) -> QueryBuilder

Keep only entities with at least one qualifying assertion whose effective trust_level >= level.

Current-state (no .as_of()): active by default, also flagged when .include_flagged() is set and/or superseded/retracted when .include_history() is set (KI-093). Under .as_of(t), .include_flagged() still applies (excludes flagged point-in-time, KI-097, not by current status) and .include_history() has nothing to add for a superseded assertion — but does for a retracted one (ADR-0049, KI-095), same as .min_confidence()'s identical carve-out above.

"Effective" (KI-047): for an assertion made under delegation (acting_as set), this is min(author.trust_level, acting_as.trust_level) — the same effective-trust formula govern/policy.py already uses to decide whether to auto-accept that same assertion, by analogy with SPEC §8.4's capability rule ("effective capability is min(author, acting_as)") — not the author's raw trust_level alone. For a non-delegated assertion it's simply the author's own trust_level.

Respects .as_of() (KI-036) for which assertion counts as qualifying, the same way .min_confidence() does. Each principal's own trust_level is always its current value: no code path ever changes a principal's trust_level after creation, so there is no historical value to reconstruct — "as of t" and "now" are the same number by construction, and the same holds for the min() this method now takes of two such principals' trust levels.

Parameters:

Name Type Description Default
level int

Minimum effective trust level, 0-10.

required

Returns:

Type Description
QueryBuilder

Self for chaining

Source code in src/ontolith/query/builder.py
def trust_at_least(self, level: int) -> "QueryBuilder":
    """Keep only entities with at least one qualifying assertion whose
    *effective* trust_level >= `level`.

    Current-state (no `.as_of()`): `active` by default, also `flagged`
    when `.include_flagged()` is set and/or `superseded`/`retracted`
    when `.include_history()` is set (KI-093). Under `.as_of(t)`,
    `.include_flagged()` still applies (excludes `flagged`
    point-in-time, KI-097, not by current status) and
    `.include_history()` has nothing to add for a `superseded`
    assertion — but does for a `retracted` one (ADR-0049, KI-095),
    same as `.min_confidence()`'s identical carve-out above.

    "Effective" (KI-047): for an assertion made under delegation
    (`acting_as` set), this is `min(author.trust_level,
    acting_as.trust_level)` — the same effective-trust formula
    `govern/policy.py` already uses to decide whether to auto-accept
    that same assertion, by analogy with SPEC §8.4's capability rule
    ("effective capability is min(author, acting_as)") — not the
    author's raw trust_level alone. For a non-delegated assertion it's
    simply the author's own trust_level.

    Respects `.as_of()` (KI-036) for which assertion counts as
    qualifying, the same way `.min_confidence()` does. Each principal's
    own trust_level is always its current value: no code path ever
    changes a principal's trust_level after creation, so there is no
    historical value to reconstruct — "as of t" and "now" are the same
    number by construction, and the same holds for the `min()` this
    method now takes of two such principals' trust levels.

    Args:
        level: Minimum effective trust level, 0-10.

    Returns:
        Self for chaining
    """
    self._trust_at_least = level
    return self

limit

limit(n: int) -> QueryBuilder

Cap the number of entities .all() returns.

Parameters:

Name Type Description Default
n int

Maximum number of results.

required

Returns:

Type Description
QueryBuilder

Self for chaining

Source code in src/ontolith/query/builder.py
def limit(self, n: int) -> "QueryBuilder":
    """Cap the number of entities `.all()` returns.

    Args:
        n: Maximum number of results.

    Returns:
        Self for chaining
    """
    self._limit = n
    return self

include_flagged

include_flagged() -> QueryBuilder

Also match .where() filters against flagged assertions (SPEC §11.2, §10.3, KI-081).

By default a .where() filter matches only active assertions, so an entity whose only matching value sits on a flagged (contradicted) assertion is not returned. This opt-in widens the match set to active + flagged. It is a match-set widener, not a result-shape change — .all() still returns list[Entity].

No effect on a query with neither a .where() filter nor a .min_confidence()/.trust_at_least() floor — nothing to widen (KI-093: those floors are evaluated even when .where() is absent, so this flag is not a no-op just because .where() wasn't called — only when nothing downstream inspects assertion status at all).

Does apply under .as_of(t), for both .where() and, since KI-093, .min_confidence()/.trust_at_least() — but not by current status: flagged is reconstructed point-in-time from the assertion_event log (KI-097), so a query pinned to a t when an assertion was disputed correctly excludes it even after the dispute has since been resolved, and a t strictly before any dispute existed still includes an otherwise-undisputed value, even though the same assertion is flagged right now.

Returns:

Type Description
QueryBuilder

Self for chaining

Source code in src/ontolith/query/builder.py
def include_flagged(self) -> "QueryBuilder":
    """Also match `.where()` filters against `flagged` assertions
    (SPEC §11.2, §10.3, KI-081).

    By default a `.where()` filter matches only `active` assertions, so
    an entity whose only matching value sits on a flagged (contradicted)
    assertion is not returned. This opt-in widens the match set to
    `active` + `flagged`. It is a *match-set* widener, not a result-shape
    change — `.all()` still returns `list[Entity]`.

    No effect on a query with neither a `.where()` filter nor a
    `.min_confidence()`/`.trust_at_least()` floor — nothing to widen
    (KI-093: those floors are evaluated even when `.where()` is absent,
    so this flag is *not* a no-op just because `.where()` wasn't
    called — only when nothing downstream inspects assertion status at
    all).

    Does apply under `.as_of(t)`, for both `.where()` and, since
    KI-093, `.min_confidence()`/`.trust_at_least()` — but not by
    current status: `flagged` is reconstructed point-in-time from the
    `assertion_event` log (KI-097), so a query pinned to a `t` when an
    assertion *was* disputed correctly excludes it even after the
    dispute has since been resolved, and a `t` strictly before any
    dispute existed still includes an otherwise-undisputed value, even
    though the same assertion is `flagged` right now.

    Returns:
        Self for chaining
    """
    self._include_flagged = True
    return self

include_history

include_history() -> QueryBuilder

Also match .where() filters against superseded and retracted assertions (SPEC §11.2, KI-081).

By default a .where() filter matches only active assertions. This opt-in widens the current-state match set to also include assertions a later write superseded or a retract() withdrew — so kb.query(Person).where(name="Ada").include_history() returns a person whose name was "Ada" even if it isn't now. Combine with .include_flagged() to widen to every status.

Like .include_flagged(), a match-set widener, not a result-shape change: .all() returns list[Entity], with no per-entity timeline attached. Also widens .min_confidence()/ .trust_at_least() (KI-093) the same way it widens .where() — see .include_flagged()'s docstring for why this means it is not a no-op just because .where() wasn't called.

Under .as_of(t), mostly no effect, but not entirely (ADR-0049, KI-095): the as_of branch never restricts matches to active in the first place — it excludes only flagged (itself gated behind .include_flagged()) — so a superseded assertion whose validity window covers t already matches there regardless of this opt-in. A retracted assertion is different: as_of(t) additionally excludes it once its own retraction event's assertion-time has passed (t is at or after when the retraction was recorded) — .include_history() opts back out of that exclusion, the one thing it does on the as_of path. True for both the .where() path and, since KI-093, the confidence/trust floors.

Returns:

Type Description
QueryBuilder

Self for chaining

Source code in src/ontolith/query/builder.py
def include_history(self) -> "QueryBuilder":
    """Also match `.where()` filters against `superseded` and
    `retracted` assertions (SPEC §11.2, KI-081).

    By default a `.where()` filter matches only `active` assertions.
    This opt-in widens the current-state match set to also include
    assertions a later write superseded or a `retract()` withdrew — so
    `kb.query(Person).where(name="Ada").include_history()` returns a
    person whose name *was* "Ada" even if it isn't now. Combine with
    `.include_flagged()` to widen to every status.

    Like `.include_flagged()`, a match-set widener, not a result-shape
    change: `.all()` returns `list[Entity]`, with no per-entity
    timeline attached. Also widens `.min_confidence()`/
    `.trust_at_least()` (KI-093) the same way it widens `.where()` —
    see `.include_flagged()`'s docstring for why this means it is
    *not* a no-op just because `.where()` wasn't called.

    Under `.as_of(t)`, mostly no effect, but not entirely (ADR-0049,
    KI-095): the `as_of` branch never restricts matches to `active` in
    the first place — it excludes only `flagged` (itself gated behind
    `.include_flagged()`) — so a `superseded` assertion whose validity
    window covers `t` already matches there regardless of this opt-in.
    A `retracted` assertion is different: `as_of(t)` additionally
    excludes it once its own retraction event's assertion-time has
    passed (`t` is at or after when the retraction was recorded) —
    `.include_history()` opts back out of *that* exclusion, the one
    thing it does on the `as_of` path. True for both the `.where()`
    path and, since KI-093, the confidence/trust floors.

    Returns:
        Self for chaining
    """
    self._include_history = True
    return self

all

all() -> list[Entity]

Execute query and return all matching entities.

Returns:

Type Description
list[Entity]

List of matching entities (may be empty)

Source code in src/ontolith/query/builder.py
def all(self) -> list[Entity]:
    """Execute query and return all matching entities.

    Returns:
        List of matching entities (may be empty)
    """
    entities = self._base_candidates()
    entities = self._apply_confidence_trust_filters(entities)
    if self._limit is not None:
        entities = entities[: self._limit]
    return entities

first

first() -> Entity | None

Execute query and return first matching entity.

Returns:

Type Description
Entity | None

First matching entity, or None if no matches

Source code in src/ontolith/query/builder.py
def first(self) -> Entity | None:
    """Execute query and return first matching entity.

    Returns:
        First matching entity, or None if no matches
    """
    results = self.all()
    return results[0] if results else None

count

count() -> int

Execute query and return count of matching entities.

Returns:

Type Description
int

Number of matching entities

Source code in src/ontolith/query/builder.py
def count(self) -> int:
    """Execute query and return count of matching entities.

    Returns:
        Number of matching entities
    """
    return len(self.all())