Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 54 additions & 2 deletions packages/gooddata-eval/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -136,7 +136,7 @@ gd-eval run \
| `--reasoning-effort LEVEL` | server default | `LOW`, `MEDIUM` or `HIGH`, sent as `options.reasoningEffort` on every chat message. Requires the `enableGenAiReasoningEffort` feature flag on the target organization — without it the server ignores the value. Applies to chat items only; `dashboard_summary` items go through the summary endpoint, which has no such option. |

**Concurrency and workspace safety.** Agentic kinds that create workspace objects
(`agentic_metric_skill`, `agentic_alert_skill`, `agentic_conversation`, `agentic_kda_skill`) always run one at a
(`agentic_metric_skill`, `agentic_alert_skill`, `agentic_conversation`, `agentic_kda_skill`, `agentic_obfuscation`) always run one at a
time whatever `--concurrency` says — a metric or alert created and dropped mid-run would otherwise be visible to
another item reading the same catalog. **That protection is for the agentic kinds only:** the single-turn
`metric_skill` and `alert_skill` kinds are still fanned out and the agent performs the same server-side writes on
Expand Down Expand Up @@ -535,7 +535,8 @@ A dataset is a folder of `.json` files, one per question:
```

Supported `test_kind` values: `visualization`, `metric_skill`, `alert_skill`,
`search_tool`, `general_question`, `guardrail`, `dashboard_summary`.
`search_tool`, `general_question`, `guardrail`, `dashboard_summary`, and the agentic kinds
(`agentic_obfuscation` is described below).

### `dashboard_summary` items

Expand Down Expand Up @@ -574,6 +575,53 @@ The `expected_output` rubric:
Each criterion is scored independently by the LLM judge, so `quality_score`
is the fraction of satisfied criteria.

### `agentic_obfuscation` items

gen-ai masks sensitive values out of the conversation it stores and the trace it exports to
Langfuse, while the model and the user's stream keep what was typed. So these items are not
graded on the answer: they plant synthetic *canaries* in one or more user turns, then read back
the stored conversation (`GET …/chat/conversations/{id}/items`) and every Langfuse trace of the
session, and decide by exact substring. No LLM decides whether a value leaked.

```json
{
"id": "obfuscation-001",
"dataset_name": "agent_obfuscation",
"test_kind": "agentic_obfuscation",
"question": ["My email is qa.canary5a1e@example.invalid. Show Total Sales by month.", "Now by quarter instead."],
"expected_output": {
"status": "enforced",
"canaries": [
{"nonce": "canary5a1e", "value": "qa.canary5a1e@example.invalid", "class": "EMAIL",
"absent_from": ["conversation_db", "langfuse_trace"], "mask_marker_present": "[EMAIL]"}
]
}
}
```

- `question` is a string, or a list of turns sent in order to one conversation (from Langfuse, an
input list). A question that is itself a JSON document must be stored in Langfuse as
`{"query": "<the JSON>"}`: Langfuse parses a JSON-looking string input into an object.
- A canary lists the sinks it must be `absent_from` and the sinks it must stay `present_in`. With
`status: known_limitation` a `present_in` value is a known gap: the item fails once it closes, so
the fixture is flipped deliberately. With `status: enforced` it guards a value that is not
sensitive: masking it fails the item as `OVER-MASKED`. `record_only_paths` names sink paths
reported but not gated.
- `expected_turn_rejected: {"status_code": 422, "reason": "DATA_OBFUSCATION_CONTENT_REJECTED"}`
expects the first turn to be refused.
- `observe` (`automation_match`, `metric_title`, `stream_markers`) reports what the chat created
or streamed, never gates it, and deletes the alert, export or metric it recognises.
- Every turn needs a non-sensitive fragment of 12+ word characters, the *anchor*: a sink read back
without every anchor is reported as blind, never as clean. `anchors` overrides the derived ones.
- Langfuse is polled for the session's traces for 60 s (`GD_EVAL_OBFUSCATION_LANGFUSE_TIMEOUT_SEC`).
When none arrives the run fails with `TRACE_NOT_FOUND`, the report's `trace_found` is false and
`obfuscation_trace_found` is 0 -- a missing trace points at the export, never counts as a pass.

A leak in any run fails the item whatever `--gate` says. Before the first item of a workspace a
preflight probe checks that both legs mask and both read-backs work, and stops every item of that
workspace when they do not. The kind needs the `LANGFUSE_*` credentials of the Langfuse project
the environment exports to: the traces are read, not only scored.

## Supported test kinds

| test_kind | What the agent must produce | Extra required |
Expand All @@ -585,6 +633,7 @@ is the fraction of satisfied criteria.
| `general_question` | Text answer judged by LLM | `[llm-judge]` |
| `guardrail` | Refusal/redirect (visualization response auto-fails) | `[llm-judge]` |
| `dashboard_summary` | Dashboard summary (via `/summary` endpoint) scored against a rubric by LLM | `[llm-judge]` |
| `agentic_obfuscation` | Canaries masked in the stored conversation and the Langfuse trace (exact match) | `LANGFUSE_*` |

## Optional extras

Expand Down Expand Up @@ -625,6 +674,9 @@ the item's own root span. On the agentic path each score is mirrored onto the ag
| `quality_score` | Fraction of strict check flags that are `True` (0.0–1.0). Shown in CLI as a percentage. |
| `value_score` | Weighted blend: 0.6 × quality + 0.2 × speed (speed = max(0, 1 − latency/60s)). |
| `latency_s` | Average per-run latency in seconds. |
| `obfuscation_pass` / `obfuscation_no_leak` | `agentic_obfuscation`: the run's verdict, and whether no canary leaked. The comment names what failed, with canary values replaced by their class. |
| `obfuscation_trace_found` | `agentic_obfuscation`: 0 when Langfuse held no trace for the conversation within the wait (`TRACE_NOT_FOUND`). |
Comment thread
myhoai marked this conversation as resolved.
| `obfuscation_record_only_hit` | `agentic_obfuscation`: 1 when a record-only path held a canary; the comment carries every record-only observation. |
| `provider_type` | Model vendor + gateway label (e.g. `ANTHROPIC`, `BEDROCK/ANTHROPIC`, `AZURE/OPENAI`). Stored in Langfuse trace metadata and tags. |

Score names carry no K; K and the gate are on the dataset-run metadata as `eval_k` and `eval_gate`.
Expand Down
19 changes: 18 additions & 1 deletion packages/gooddata-eval/src/gooddata_eval/cli/agentic_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@
from gooddata_eval.core.agentic.guardrail import evaluate_agentic_guardrail
from gooddata_eval.core.agentic.kda_skill import evaluate_agentic_kda_skill
from gooddata_eval.core.agentic.metric_skill import evaluate_agentic_metric_skill
from gooddata_eval.core.agentic.obfuscation import evaluate_agentic_obfuscation
from gooddata_eval.core.agentic.search_tool import evaluate_agentic_search_tool
from gooddata_eval.core.agentic.visualization import evaluate_agentic_visualization
from gooddata_eval.core.agentic.what_if import evaluate_agentic_what_if
Expand Down Expand Up @@ -49,6 +50,7 @@ class _LfKw(TypedDict, total=False):
"agentic_conversation",
"agentic_kda_skill",
"agentic_what_if",
"agentic_obfuscation",
}
)

Expand Down Expand Up @@ -88,7 +90,9 @@ class _LfKw(TypedDict, total=False):
# metric skill. agentic_kda_skill is here on suspicion rather than proof: it triggers
# create_key_driver_analysis with no cleanup, and while the evaluator only ever reads that
# call's ARGUMENTS -- never a created object id -- whether the platform persists anything is
# unverified. Move it to the allowlist once someone confirms it does not.
# unverified. Move it to the allowlist once someone confirms it does not. agentic_obfuscation
# items may ask for an alert, a scheduled export or a metric; it deletes what it recognises,
# but that is cleanup, not read-only.
#
# agentic_dashboard_skill is absent by default rather than by evidence: gen-ai holds the draft and
# any chart it authors in conversation state and writes neither until a user saves from the UI, so
Expand Down Expand Up @@ -278,6 +282,19 @@ def _dispatch_agentic(
agent_id=agent_id,
**lf_kw,
)
elif kind == "agentic_obfuscation":
return evaluate_agentic_obfuscation(
host=host,
token=token,
workspace_id=workspace_id,
question=item.question,
expected_output=eo if isinstance(eo, dict) else {},
k=k,
turns=item.turns,
gate=gate,
agent_id=agent_id,
**lf_kw,
)
elif kind == "agentic_conversation":
fixture_data = eo.get("fixture") or eo if isinstance(eo, dict) else {}
return evaluate_agentic_conversation(
Expand Down
13 changes: 12 additions & 1 deletion packages/gooddata-eval/src/gooddata_eval/cli/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,17 @@ def close(self) -> None:
backend.close()


def _positive_int(value: str) -> int:
"""An argparse type for counts that must be at least 1."""
try:
number = int(value)
except ValueError as exc:
raise argparse.ArgumentTypeError(f"not an integer: {value!r}") from exc
if number < 1:
raise argparse.ArgumentTypeError(f"must be at least 1, got {number}")
return number


def _build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(prog="gd-eval", description="Evaluate the GoodData AI agent.")
sub = parser.add_subparsers(dest="command", required=True)
Expand Down Expand Up @@ -102,7 +113,7 @@ def _build_parser() -> argparse.ArgumentParser:
"Default: workspace's current active model."
),
)
run.add_argument("--runs", type=int, default=2, help="Independent runs per item. Default 2.")
run.add_argument("--runs", type=_positive_int, default=2, help="Independent runs per item. Default 2.")
run.add_argument(
"--gate",
choices=get_args(EvalGate),
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,223 @@
# (C) 2026 GoodData Corporation. All rights reserved.
"""Deterministic leak verdict for the agentic_obfuscation kind.

Pure functions only: the caller collects the sinks, this module decides. A sink is any
JSON-like document -- the stored conversation read back over the API, or every Langfuse
trace of the conversation's session. A canary is found by exact substring over every string
leaf, so the verdict never depends on an LLM. JSON serialised inside a string (Langfuse
keeps span input that way) is decoded and walked too.

A sink that cannot be seen must not pass an absence check. Every turn therefore carries an
anchor, a non-sensitive fragment of the question, and a sink that does not show every
anchor is reported as blind instead of clean.
"""

from __future__ import annotations

import json
import re
from collections.abc import Iterator, Mapping, Sequence
from dataclasses import dataclass, field
from typing import Any

SINK_CONVERSATION_DB = "conversation_db"
SINK_LANGFUSE_TRACE = "langfuse_trace"
SINKS = (SINK_CONVERSATION_DB, SINK_LANGFUSE_TRACE)

# Classes whose value may be written with separators the detector ignores, so a leak of
# "4916 3385 0608 2832" must still match the compact nonce "4916338506082832".
_DIGIT_CLASSES = frozenset({"CREDIT_CARD", "IBAN"})
_SEPARATORS = re.compile(r"[\s\-]")
# Placeholder the AI-on-data trace redaction writes over Langfuse input and output. Behind it
# the Langfuse leg shows nothing, so an absence check there would prove nothing.
AI_ON_DATA_REDACTION = "<redacted: AI-on-data enabled>"
_ANCHOR_RUN = re.compile(r"[^\W_](?:[\w ]*[^\W_])?")
_ANCHOR_MIN = 12
_ANCHOR_MAX = 40


@dataclass(frozen=True)
class Occurrence:
sink: str
path: str
needle: str


@dataclass
class CaseVerdict:
failures: list[str] = field(default_factory=list)
notes: list[str] = field(default_factory=list)

@property
def passed(self) -> bool:
return not self.failures


def iter_strings(value: Any, path: str = "$") -> Iterator[tuple[str, str]]:
"""Yield ``(path, text)`` for every string or number leaf and every key, decoding JSON held in a string.

A string that parses as a JSON document is walked instead of being yielded whole, so a
value spelled with escapes (``\\u0040``) is seen decoded and reported once. Numbers are
leaves too: a numeric secret in a pasted JSON document decodes to an int. Keys are text a
user can type as well -- a pasted config may key on an email -- so each is yielded at the
path of its object with a ``<key>`` suffix.
"""
if isinstance(value, bool):
return
if isinstance(value, int | float):
yield path, str(value)
return
if isinstance(value, str):
if value.lstrip()[:1] in ("{", "["):
try:
decoded = json.loads(value)
except ValueError:
pass
else:
yield from iter_strings(decoded, f"{path}<json>")
return
yield path, value
elif isinstance(value, Mapping):
for key, child in value.items():
if isinstance(key, str):
yield f"{path}.{key}<key>", key
yield from iter_strings(child, f"{path}.{key}")
elif isinstance(value, Sequence) and not isinstance(value, bytes | bytearray):
for index, child in enumerate(value):
yield from iter_strings(child, f"{path}[{index}]")


def canary_needles(canary: Mapping[str, Any]) -> list[str]:
"""Every spelling of the canary a leak may take: nonce, value and its declared variants."""
keys = ("nonce", "value", "unescaped_value", "decoded_value", "compact_value")
needles = [canary[key] for key in keys if isinstance(canary.get(key), str) and canary[key]]
return list(dict.fromkeys(needles))


def find_canary(canary: Mapping[str, Any], sink: str, document: Any) -> list[Occurrence]:
needles = canary_needles(canary)
compact = {_SEPARATORS.sub("", n) for n in needles} if canary.get("class") in _DIGIT_CLASSES else set()
found: list[Occurrence] = []
for path, text in iter_strings(document):
hit = next((n for n in needles if n in text), None)
if hit is None and compact:
squeezed = _SEPARATORS.sub("", text)
hit = next((n for n in compact if n in squeezed), None)
if hit is not None:
found.append(Occurrence(sink, path, hit))
return found


def contains(document: Any, needle: str) -> bool:
return any(needle in text for _, text in iter_strings(document))


def derive_anchor(question: str, canaries: Sequence[Mapping[str, Any]]) -> str | None:
"""The longest plain-word run of the question once every canary spelling is cut out.

Word characters and spaces only, so escaping, JSON quoting and masking next to a canary
cannot alter it between what was sent and what a sink stores.
"""
fragments = [question]
# Longest spelling first: cutting the nonce out of "john\\.canary2d8e@..." first would
# leave "john\\." behind, and the anchor would then end in text the mask replaces.
needles = sorted({n for canary in canaries for n in canary_needles(canary)}, key=len, reverse=True)
for needle in needles:
fragments = [piece for fragment in fragments for piece in fragment.split(needle)]
runs = [match.group(0) for fragment in fragments for match in _ANCHOR_RUN.finditer(fragment)]
best = max(runs, key=len, default="")
if len(best) < _ANCHOR_MIN:
return None
return best[:_ANCHOR_MAX].rstrip()


def _declared_sinks(canary: Mapping[str, Any], key: str) -> list[str]:
sinks = canary.get(key) or []
unknown = [s for s in sinks if s not in SINKS]
if unknown:
raise ValueError(f"canary {canary.get('nonce')!r} names unknown sink(s) {unknown} in {key}")
return list(sinks)


def _format(occurrences: Sequence[Occurrence], limit: int = 5) -> str:
shown = ", ".join(f"{o.path} ({o.needle!r})" for o in occurrences[:limit])
more = len(occurrences) - limit
return shown + (f" and {more} more" if more > 0 else "")


def evaluate_case(
expected: Mapping[str, Any],
sinks: Mapping[str, Any],
anchors: Sequence[str],
*,
turn_rejected: bool = False,
) -> CaseVerdict:
"""Decide one item from the collected sinks.

``sinks`` maps a sink name to its document, or to ``None`` when it was not collected
(only allowed for the Langfuse leg of a rejected turn, which may export nothing).
"""
verdict = CaseVerdict()
status = expected.get("status", "enforced")
visible: dict[str, bool] = {}

for sink in SINKS:
document = sinks.get(sink)
if document is None:
visible[sink] = False
if not turn_rejected:
verdict.failures.append(f"{sink}: not collected, so no absence claim can be made")
continue
if sink == SINK_LANGFUSE_TRACE and contains(document, AI_ON_DATA_REDACTION):
verdict.failures.append(
f"{sink}: trace input/output is replaced by {AI_ON_DATA_REDACTION!r} (enableAiOnData with "
"enableGenAiTraceRedaction), so the obfuscation leg cannot be observed here"
)
visible[sink] = False
continue
missing = [anchor for anchor in anchors if not contains(document, anchor)]
if missing and not turn_rejected:
verdict.failures.append(f"{sink}: blind -- anchor(s) {missing} not found, the read-back is incomplete")
visible[sink] = False
continue
visible[sink] = True

for canary in expected.get("canaries", []):
label = f"{canary.get('class')} {canary.get('nonce')!r}"
for sink in _declared_sinks(canary, "absent_from"):
# A blind sink still convicts: a canary it does show is a leak all the same.
if sinks.get(sink) is None:
continue
occurrences = find_canary(canary, sink, sinks[sink])
# Paths the item declares as not yet decided (e.g. conversation state the SC does
# not rule on) are reported, never gated, until a decision turns them into leaks.
record_only = [re.compile(p) for p in canary.get("record_only_paths") or []]
recorded = [o for o in occurrences if any(r.search(o.path) for r in record_only)]
gated = [o for o in occurrences if o not in recorded]
if gated:
verdict.failures.append(f"LEAK {label} in {sink}: {_format(gated)}")
if recorded:
verdict.notes.append(f"RECORDED {label} in {sink} (record-only path, not gated): {_format(recorded)}")
marker = canary.get("mask_marker_present")
if marker and not turn_rejected and visible.get(sink) and not contains(sinks[sink], marker):
verdict.failures.append(f"{label}: mask marker {marker!r} missing from {sink}")
for sink in _declared_sinks(canary, "present_in"):
if not visible.get(sink):
continue
if not find_canary(canary, sink, sinks[sink]):
if status == "known_limitation":
flip = (expected.get("known_limitation") or {}).get("flip_when", "")
verdict.failures.append(
f"{label} expected in {sink} ({status}) but is now masked -- the limitation no longer "
f"reproduces, update the fixture deliberately. flip_when: {flip}"
)
else:
# An enforced present_in guards a value that is not sensitive: masking it is
# the defect (a false positive), not progress.
verdict.failures.append(
f"OVER-MASKED {label} in {sink}: expected unchanged but it was masked (false positive)"
)

if turn_rejected and sinks.get(SINK_LANGFUSE_TRACE) is None:
verdict.notes.append("rejected turn exported no Langfuse trace; only the database leg was asserted")
return verdict
Loading
Loading