grithdocs

Performance & tuning

Measuring grith overhead and squeezing latency out of the pipeline.

grith is designed for low overhead — a typical call exits the pipeline in 8–12ms wall-clock. This page covers measuring that and the knobs you can turn when it matters.

Targets

MetricTarget
Per-call pipeline latency (median)8–12ms
Per-call pipeline latency (p99)15ms
Supervisor overhead per syscall (intercepted)~50–100µs
Supervisor overhead per syscall (uninteresting, seccomp-fast-path)under 1µs
Daemon idle RSS~30MB
Daemon busy RSS80–150MB

Measuring

Built-in: GET /proxy/status/full (IPC) returns per-filter timing histograms.

For p99 over time:

grith audit export --since 1d --format json |
  jq '[.[] | .filter_timing_ms] | sort_by(.) | .[length*99/100]'

For per-filter breakdown, the trace-mode export captures filter timing per call:

grith exec --trace-syscalls-jsonl /tmp/trace.jsonl -- claude
# ... use the agent ...
grith profile audit --trace /tmp/trace.jsonl --report-timing

Common bottlenecks

1. Secret scanner on large payloads

Symptom: p99 latency exceeds 20ms on calls with large content.

The scanner compiles all patterns into a single automaton, so cost is linear in payload size, not pattern count - and there is no scan-length cap: large outbound payloads are scanned in full. If a workload routinely pushes multi-MB payloads through supervised network sends, expect scan time proportional to size. The lever is on the workload side (smaller payloads through supervised calls); trimming patterns from config/filters/secrets.toml does not meaningfully change scan time.

2. Behavioural baseline computation

Symptom: latency spikes on certain calls after a long session.

The behavioural filter compares each call against a per-session rolling baseline. If profiling shows it hot in your workload, it is a Phase 3 filter and can be disabled outright:

[proxy.filters.behavioural]
enabled = false

Trade-off: you lose the out-of-pattern signal (sudden bursts, unusual destinations) for the session.

3. Audit log fsync

Symptom: high disk wait, even with fast SSD.

Default: fsync per audit event. For lower disk IO:

[general]
audit_fsync = "batch"        # "per-event" (default), "batch", "none"
audit_fsync_batch_ms = 100

none is unsafe; events can be lost on crash. batch is the right choice for high-throughput workloads.

4. Reputation table size

Symptom: daemon RSS climbs to 500MB+.

The reputation table accumulates without bound by default. For very long- running daemons:

[reputation]
max_entries = 50000          # cap; oldest decay first
gc_interval_seconds = 3600

Noise reduction knobs

These reduce syscall count (not latency per call), often the bigger win:

[supervisor.noise_reduction]
ignore_read_only = true              # skip filter pipeline for reads on open fds
batch_rapid_reads = true             # coalesce rapid reads in same fd
batch_window_ms = 50                 # batching window

Both default-on. Disable only for forensic recording.

When grith is "slow"

If an agent feels noticeably slow under grith:

  1. Check GET /proxy/status/full — is the filter pipeline actually the bottleneck, or is it the model latency?
  2. Check the audit log — is the agent making 10× the syscalls it would without grith? Filters can't help with sheer syscall count.
  3. Try noise-reduction off vs on — sometimes coalescing reads hides pathological agent behaviour worth seeing.

The cost model of grith is per-syscall, not per-second. Idle agents pay nothing. Busy agents pay proportional to their syscall rate.

See also

Last updated: 2026-05-14Edit this page on GitHub →