Fuzzing Guide
Observal is fuzzed with Atheris and is being submitted to Google OSS-Fuzz for continuous fuzzing; the upstream project is not live yet. The fuzz targets live here so they are versioned alongside the code they exercise; OSS-Fuzz only stores the three-file project configuration, which is mirrored under fuzz/oss-fuzz/.
Targets
session_jsonl_fuzzer
Harness session transcripts: ingest_classify on the write path, parse_raw_events on the read path
session_structure_fuzzer
Same pipeline, driven by Hypothesis-generated transcript records instead of raw bytes
secrets_redactor_fuzzer
services.secrets_redactor.redact_secrets, applied to every line and preview before storage
support_redaction_fuzzer
observal_cli.support.redaction, the single chokepoint for observal doctor support bundle
Every target runs entirely in-process. None of them opens a socket, reads credentials, touches PostgreSQL, ClickHouse or Redis, or depends on the clock, so a crash reproduces from its input alone.
session_jsonl_fuzzer reads the first input byte as a harness selector (byte % len(HARNESSES)) and treats the rest as the transcript, which is why each seed corpus file starts with a single digit and is otherwise plain JSONL.
Running a target locally
pip install atheris hypothesis
python3 fuzz/session_jsonl_fuzzer.py -atheris_runs=100000Point libFuzzer at the seed corpus and dictionary to start from useful inputs and keep anything new it discovers:
python3 fuzz/session_jsonl_fuzzer.py \
-dict=fuzz/dictionaries/session_jsonl_fuzzer.dict \
/tmp/observal-corpus fuzz/corpus/session_jsonl_fuzzersession_structure_fuzzer is a polyglot. Run it under pytest to replay and shrink any example Hypothesis has already recorded:
pytest fuzz/session_structure_fuzzer.pymake test-fuzz runs a short campaign over every target and is the quickest way to confirm the harnesses still build after a parser change.
Reproducing a crash
A crashing input is a plain file. OSS-Fuzz attaches one to every report; locally libFuzzer writes crash-<sha1> into the working directory.
Atheris prints the Python traceback and exits non-zero. For an OSS-Fuzz report, infra/helper.py reproduce observal <target> <testcase> runs the same input inside the container the report came from.
Adding a target
Add
fuzz/<name>_fuzzer.py.build.shglobs*_fuzzer.py, so no build change is needed.Import the code under test inside
with atheris.instrument_imports():so Atheris can add coverage instrumentation as the modules load. Call_paths.add_source_roots()first if the target touchesobserval-server.Define
TestOneInput(data: bytes), wire it up inmain(), and keep the harness thin -- decode the input, call the boundary, assert the contract. Expected rejections (a decode error on malformed input) should return; every other exception is a finding.Bound the input size. Slow inputs are reported as timeouts rather than as bugs.
Declare any third-party import in the
fuzzdependency group inpyproject.tomland commit the refresheduv.lock. OSS-Fuzz installs that group into the base image's system interpreter, which is the only environment PyInstaller resolves imports against; a package that is merely importable in a local virtualenv is silently omitted from the built target.Add a seed corpus under
corpus/<name>_fuzzer/and a dictionary atdictionaries/<name>_fuzzer.dict. Both are optional and both are picked up automatically.Corpus files are raw fuzzer input, not source. Nothing may rewrite them.
scripts/update_spdx_copyright.pyskipsfuzz/corpus/**, andREUSE.tomlcarries their licensing instead.session_jsonl_fuzzerreads its first byte, so a stray header would repoint the whole corpus at one parser;tests/test_fuzz_targets.pyasserts every parser still has a seed.Keep seed corpora free of anything that looks like a credential. The repository runs Gitleaks and a pre-commit secret scan; put distinctive vendor prefixes in the dictionary, where they are too short to match a scanner rule, and use zero-entropy placeholders in corpus files.
Prefer extending an existing target over adding a near-duplicate. A new target should reach a boundary the current four do not.
Maintaining the OSS-Fuzz project
fuzz/oss-fuzz/ mirrors projects/observal/ in google/oss-fuzz. Edit the files under fuzz/oss-fuzz/, then copy them upstream in the same change so the two never drift.
Validate a change against the real builder before opening the upstream pull request:
Passing the checkout path to build_fuzzers replaces the git clone in the Dockerfile with the local tree, which is how you test targets before they land on main. Two side effects to know about: the build writes *.pkg.spec into the checkout, and the coverage build prepends a coverage stub to each target source in place. Only *.pkg.spec and coverage_wrapper.py are gitignored. Target sources are tracked, so inspect git diff fuzz/ and restore any modified targets before committing.
Once the first hosted build succeeds, add the status badge to the README:
Last updated
Was this helpful?