Skip to content

Security Architecture

This document describes the security architecture of Diogenes for security professionals evaluating the system's threat model, cryptographic design, and trust boundaries.


Design Principles

  1. Subscriber private keys never leave the client. All subscriber signing is performed client-side: keys are generated in the browser with crypto.subtle.generateKey (static/js/keys_register.js) and the server only ever receives public keys and signatures. That client-side generation is the SC-01 control.

    PrivateKeyRejectionMiddleware is defence in depth, not the SC-01 control. It matches PEM markers in POST/PUT/PATCH request bodies only, and is trivially evaded by base64 or any other encoding. It catches an accidental paste; it stops nothing deliberate.

    One carve-out, by design: the log operator's own signing key is server-held. It is loaded from operator.pem via DIOGENES_OPERATOR_KEY_PATH (AWS Secrets Manager diogenes/production/operator-key in production) through deps.get_operator_key_provider() — see src/diogenes/server/api/endorsements.py. SC-01 scopes subscriber keys; the operator key is outside that scope and is protected by secret storage and IAM instead.

  2. Content never reaches the server. Documents are hashed locally. Only the SHA-256 digest is transmitted for verification. This preserves document confidentiality.

  3. Append-only log. The transparency log is append-only and hash-chained. Once an entry is committed, it cannot be altered without breaking the chain. Chain integrity can be verified by any party.

  4. No central authority. Trust is established through a decentralized web of endorsements. There is no certificate authority whose compromise would undermine the entire system.

  5. Defense in depth. Multiple security layers operate independently: cryptographic verification, key status checking, trust assessment, rate limiting, password protection, and temporal anchoring.


Cryptographic Primitives

Primitive Usage Notes
Ed25519 Key pairs, attestation signatures Recommended algorithm. Edwards-curve Digital Signature Algorithm.
ECDSA P-256 Key pairs, browser signing Compatible with Web Crypto API. SHA-256 hash.
RSA-2048 Key pairs (legacy) Supported for backward compatibility.
SHA-256 Content hashing, fingerprints, log chain Used throughout for hash computation.
Argon2id Password hashing Key passwords hashed with Argon2id before storage.
HMAC-SHA256 JWT signing JWTs signed with operator key for session management.

Trust Boundaries

graph TD
    subgraph Client["Client (Browser / CLI)"]
        KG["Key Generation
(private)"] SG["Signing
(private)"] CH["Content Hashing
(document stays local)"] end subgraph Server["Server (API)"] SV["Signature Verification"] KR["Key Registry
(public keys)"] TL["Transparency Log
(append-only)"] RL["Rate Limiting"] SH["Security Headers
(CSP, HSTS)"] PKR["Private Key Rejection"] end subgraph DB["Database (PostgreSQL)"] DATA["Keys, Log Entries,
Endorsements, Auth Challenges"] end subgraph BTC["Bitcoin (via OpenTimestamps)"] TA["Temporal anchoring
of log entries"] end Client -- "HTTPS
(public keys, signatures, hashes only)" --> Server Server --> DB DB --> BTC

Authentication Model

Challenge-Response

Mutating operations (key succession, revocation, endorsement operations) require proof of key possession via a challenge-response protocol:

  1. The server issues a random nonce bound to a specific fingerprint.
  2. The client signs the nonce with their private key.
  3. The server verifies the signature against the registered public key.

Challenges expire (configurable TTL) and are single-use. Replay attacks are prevented by challenge consumption.

JWT Sessions

For browser-based portal sessions, a challenge-response exchange produces a JWT with:

  • 30-minute TTL (configurable)
  • 8-hour maximum session duration
  • Fingerprint binding (JWT is valid only for the authenticated key)
  • Optional recovery-only flag (limits operations available to recovered keys)

Password Protection

Keys can optionally register a password (Argon2id-hashed). When set, the password is required as a second factor for attestation signing and other sensitive operations.

Credential Binding

A session's identity is resolved purely by the credential row matching the JWT fingerprint. Binding a fingerprint to an account is therefore the identity grant itself, not a bookkeeping step — whoever can write that row can log in as that account.

Two consequences follow, and both are enforced rather than assumed:

  • There is one binding path. POST /api/v1/credentials/enroll requires a signed challenge, an active non-revoked issuing credential, a verified chain signature, and hardware backing on both sides. The self-service /users/{user_id}/credentials routes that did the same job with no checks at all were deleted in #502 (C-2) rather than authenticated, because a hardened path and an unhardened path doing the same job is the shape that produced this finding in the first place.
  • Binding requires proof of the key, not just of the account. CredentialService.enroll_credential takes a required proof of possession: either a signature by the key being bound, over canonical bytes naming both the account and the fingerprint, or an assertion that a WebAuthn ceremony has just proven possession for that same fingerprint. Deleting the routes alone would not have closed the finding — the next caller of the service would reintroduce it by omission.

The proof binds one (account, key) pair, so a proof captured for one account does not verify for another and a proof made with one key does not verify for a different one. It carries no nonce, deliberately: replaying it re-asserts the identical binding, which is idempotent under the unique constraint on credentials.key_fingerprint.

Known limit. Registration-time proof of possession is a separate gap: POST /api/v1/keys/register does not prove the registrant holds the key it registers (M-16, tracked in #518). Until that lands, the enrollment proof above shows control of the registered public key, which is as strong as registration made it.


Threat Model

Threats Addressed

Threat Mitigation
Server compromise Private keys never reach the server. Existing signatures remain valid. The append-only log detects tampering via hash chain verification.
Log tampering Hash chain integrity verification. Temporal anchoring to Bitcoin provides external proof of log state. Merkle tree heads enable efficient audit.
Key theft Password protection adds a second factor. Key revocation is immediate and logged. Endorser revocation alerts notify the trust network.
Sybil attacks Endorsement capacity scaling, activation delays, privilege thresholds, and over-capacity discounting prevent trust network gaming.
Replay attacks Challenge-response nonces are single-use and time-limited. JWTs have short TTLs.
Content leakage Documents are never uploaded. Only hashes are transmitted.
Rate-based abuse Per-IP rate limiting with resource-specific thresholds (key registrations, endorsements, attestations, general queries).

Threats Not Addressed (Out of Scope)

Threat Notes
Client-side key theft Diogenes does not manage private key storage. Key holders are responsible for securing their private keys.
Network-level attacks HTTPS is enforced in production (HTTPS redirect middleware), but network security is the deployment operator's responsibility.
Quantum computing Current algorithms (Ed25519, ECDSA, RSA) are not post-quantum resistant. Algorithm agility in the protocol allows future migration.

Security Headers

The SecurityHeadersMiddleware sets the following headers on all responses:

Header Value Purpose
X-Content-Type-Options nosniff Prevent MIME type sniffing
X-Frame-Options DENY Prevent clickjacking
Content-Security-Policy see below Prevent XSS and injection
Strict-Transport-Security max-age=63072000; includeSubDomains (HTTPS only) Enforce HTTPS

X-XSS-Protection is not set. The header is retired in every current browser and its legacy auditor was itself an injection vector; the row that used to claim it here was simply wrong.

Content-Security-Policy

The application policy — every route that touches session JWTs, key material or user content — is emitted verbatim as:

default-src 'self';
script-src 'self' 'nonce-<per-request>' 'wasm-unsafe-eval';
style-src 'self' 'unsafe-inline';
font-src 'self';
img-src 'self' data:;
connect-src 'self' https://api.iconify.design https://api.simplesvg.com https://api.unisvg.com;
base-uri 'self';
form-action 'self';
frame-ancestors 'none';
object-src 'none'

The nonce is minted per request by SecurityHeadersMiddleware and rendered into every inline <script> through a Jinja context processor. Because a nonce is present, browsers ignore 'unsafe-inline' in script-src entirely — so appending it back is not a rollback lever; reverting the policy commit is. 'wasm-unsafe-eval' is retained for the vendored hash-wasm argon2id derivation used by the signing PIN (#256); it permits WebAssembly compilation, not JavaScript eval.

Since inline handler attributes cannot carry a nonce, no template contains an on*= attribute: behaviour binds through data-on-* attributes and the delegated dispatcher in static/js/actions.js. That is enforced structurally by tests/test_template_csp_gate.py, which also forbids an inline <script> in any components/ template — a component may be returned standalone as an htmx fragment, on a different request with a different nonce.

Documented deviation. Three prefixes carry narrower policies that still allow script-src 'unsafe-inline' and a CDN origin:

Prefix Why
/api/docs, /api/redoc FastAPI generates the page, including an inline init script and cdn.jsdelivr.net asset URLs. Noncing it requires replacing docs_url/redoc_url with custom routes.
/docs MkDocs build output. Its inline scripts change on every docs edit, the tree is not in the repo, and the Material theme pulls unpkg.com (mermaid), Google Fonts, and api.github.com for the repository facts in its header.

All three are same-origin with the application, so the mitigation genuinely does not extend to them. Their content is generated from repo-controlled sources — our OpenAPI schema and our Markdown — and reflects no request input, so there is no known injection path into them. Prefix matching is exact-or-slash-prefixed, so a route such as /docsomething receives the strict application policy. Closing the deviation means vendoring swagger-ui, redoc and mermaid; the constants in middleware.py name that as their deletion condition.

Headers are defence in depth, not the encoding control. nosniff does nothing for a response that already declares text/html.

Output encoding

HTML is written by exactly one thing: the Jinja environment, which autoescapes every .html template. Response bodies constructed in Python must therefore be string literals — a body built from an f-string, a variable, or %/format bypasses the template engine and with it the only encoding control.

That rule is enforced by tests/test_html_response_literal_gate.py (#506 / H-5), which also asserts per template that autoescape is genuinely on and that no template reaches for |safe, Markup(, or {% autoescape false %}. Exception text reaches the user only through diogenes.server.web._errors.safe_error_detail, which names the field and the reason but never echoes the submitted value — pydantic's default str(exc) repeats the caller's input verbatim in input_value=.


Rate Limiting

Rate limits are enforced by one pure-ASGI middleware that runs ahead of routing, authentication and every dependency. Which routes are metered, and against which budget, is data — server/rate_limit_policy.BUCKETS is the single source of truth, and a test walks app.openapi() to fail the build if any published route resolves to no bucket.

Coverage is a default with named exceptions, not a list. Before #511 the middleware returned early for everything outside /api/ and dispatched inside /api/ from a hard-coded if/elif over five paths, so 52 of the 57 API write routes and all 53 web routes were unmetered (finding M-10).

Shared marks a bucket whose count is authoritative across workers, held in rate_limit_counters rather than in process memory. Before #511 nothing was: docker-entrypoint.sh runs uvicorn --workers 4, so a configured limit of N was really 4N per task and every deploy reset it (finding M-7).

Bucket Default limit Shared Purpose
API queries (api-query) 100/minute per IP no GET/HEAD/OPTIONS tier default under /api/
API writes (api-write) 120/minute per IP no tier default for every other API method — the 52 routes nobody enumerated
Key registrations (key-reg) 10/hour per IP yes Limit Sybil key creation
Endorsement offers (endorsement-offer) 20/day per IP yes Limit trust network gaming
Attestation events (attestation) 50/hour per IP yes Prevent log flooding (/documents/sign shares it, so the alias is not a bypass)
Log head anchoring (log-anchor) 12/hour per IP yes Bound anonymous flood and OTS calendar amplification
Account recovery (users-recover) 10/hour per IP yes Bound guessing against an unauthenticated account-reset primitive
Verification (verify) 30/minute per IP no The seven unauthenticated verification POSTs plus the two portal forms: pure CPU (DSSE checks), no database
Claim reveal (claim-reveal) 30/hour per IP yes Bound the guessing rate against the claim commitment (does not close M-15 — the commitment is still unsalted)
Web pages (web-page) 240/minute per IP no GET/HEAD/OPTIONS tier default outside /api/
Web writes (web-write) 60/minute per IP no tier default for every other portal method
PIN proof checks (pin-check) 5 per 15 min per fingerprint no — already durable Bound signing-PIN guessing; a pre-filter in front of the DB-backed signing_pins lockout
Authentication failures 5 per 15 min per fingerprint durable Survives restart and is shared across workers (auth_lockouts)
Recovery failures 5 per 15 min per account durable Survives restart and is shared across workers (auth_lockouts)

Exempt, deliberately: GET /api/v1/health (a 429 there fails the ALB target group's matcher = "200" and cycles the ECS task), /static/* and /docs/* (StaticFiles mounts — no database, no application CPU, and one docs page pulls dozens of assets). First match wins and buckets never stack, so a request is charged to exactly one budget and Retry-After stays meaningful; the consequence is that a client's total ceiling is the sum of the buckets it touches.

The middleware runs ahead of the auth dependencies, so the log-anchor and account-recovery buckets also cap unauthenticated floods against those routes.

Shared limiter state (#511)

Shared buckets consult a Postgres sliding-window counter after the local sliding-window log allows. The local tier keeps an already-refused client off the database and can never produce a false 429, because its limit equals the global limit. A denied request does not increment the shared counter, so a flood cannot inflate it past the limit. Retry-After in the shared tier is computed from the sliding-window estimate rather than the window edge.

Because this middleware is outermost, a store call acquires a pooled connection before the handler acquires its own. Three guards bound that: an in-flight cap of two per worker (at most 2 of 10 pooled connections), a 250 ms asyncio.wait_for around acquire and execute, and a breaker that opens for 60 seconds after five consecutive errors, timeouts or sheds.

The volumetric buckets (api-query, api-write, verify, web-page, web-write) are deliberately not shared: making them globally exact costs a database statement per request, which is itself the amplification a DoS control exists to prevent. Their boundary stays the task's concurrency and the ALB. This is the recorded residual of M-7 and the hand-off to #431, where a Redis store behind the same RateLimitStore protocol would make per-request coordination cheap enough to extend sharing.

DIOGENES_RATE_LIMIT_SHARED_STORE=none restores exact pre-#511 per-process limiting. Every limit is likewise an environment variable on the existing ECS task definition, so retuning or disabling a bucket needs no image build and no terraform apply.

Request body cap (#511)

PrivateKeyRejectionMiddleware buffers the whole body of every POST, PUT and PATCH ahead of routing and authentication, and did so with no size check (finding M-9) — an unauthenticated memory-exhaustion primitive against a 1 vCPU / 2048 MiB task shared by four workers. It now refuses an over-cap Content-Length without a single receive(), and aborts a streamed body the moment the running total crosses DIOGENES_MAX_REQUEST_BODY_BYTES (default 2 MiB, derived from the 1 MiB maximum payload plus DSSE envelope headroom) rather than draining it. The 413 carries Connection: close; the residual is that a client that already streamed megabytes may see a reset before it reads the response, and draining politely would be the denial of service.

Durable authentication and recovery lockout (#510)

The last two rows are not middleware buckets. They are rows in auth_lockouts, keyed (scope, subject) — auth on a key fingerprint, recovery on a user id — and they mirror the counter columns signing_pins already carries. Persisting them in the database rather than in process memory is the point: an in-memory counter resets on every deploy and each worker keeps its own, so five workers mean five budgets. A locked subject is answered 429 with Retry-After, and the window self-expires — no operator unlock exists or is needed.

The brake sits inside the two AuthService primitives, not on the routes. The argon2 password is the genuinely brute-forceable credential (it is human-chosen), and it is adjudicated at five places: POST /auth/login, POST /auth/password, POST /api/v1/attestations, and — through AuthService.require_auth — the key and endorsement routes. Guarding the primitive covers all five and any sixth added later. The counter is cleared by the routes instead, at the point a complete authentication succeeds: a password brute-forcer supplies a valid signature on every attempt, so a counter reset on signature success would oscillate and never arm.

POST /api/v1/auth/challenge is deliberately not guarded. It is the shared proof-minting primitive for every challenge-gated route and for the recovery-eligibility path, and fingerprints are public (transparency log, key registry), so guarding it would let six requests deny a victim every authenticated operation and their recovery path for 15 minutes, renewably, for any key in the system. It would also buy nothing: verify_challenge marks a challenge used with a flush that the request transaction rolls back when the handler raises, so a single challenge already funds unlimited guesses. The brake belongs on the guess.

Four trade-offs are accepted knowingly. Lockout is itself a denial of service: anyone who knows a fingerprint can lock it out of login, password change, signing, key operations and endorsements for 15 minutes. That is bounded by the self-expiring window, by the challenge route staying open so the victim's recovery path is never denied, and by the counter arming only on genuine credential guesses — never on a missing, expired, used or wrong-purpose challenge, nor on a "password required" refusal. And recovery answers 429 where an unknown pseudonym answers 401; that is not a meaningful enumeration oracle, because pseudonyms are published in the transparency log and the public key registry. Failures are recorded only for accounts that exist, so a flood of random pseudonyms cannot mint unbounded rows — the per-IP recovery bucket bounds that instead.

The third is fail-open on a database error. All three halves of the brake degrade to silence rather than to a refusal: the check returns "not locked", and the failure count and the post-success clear are logged and swallowed. So for as long as a database error persists, the durable brake on login, password change, every route behind AuthService.require_auth, and account recovery is inert, and a log line is the only evidence. The alternative — failing closed — is a total authentication and recovery outage on the same fault, which is why the trade is taken; docker-entrypoint.sh runs alembic upgrade head before uvicorn starts, so a task never serves traffic against a schema missing auth_lockouts, and the per-IP middleware buckets in the table above keep running independently of the database, so an unauthenticated flood is still bounded while the durable half is dark. The degradation is deliberately narrow: only a SQLAlchemyError is swallowed, so a programming error in the control fails loudly in development and CI instead of going quietly dark in production. Both DB-touching halves that run on the caller's session — the check's read and the clear's delete — do so inside a SAVEPOINT, because a failed statement aborts the enclosing Postgres transaction: without it a swallowed lockout error would either resurface as an unrelated 500 on the request's next statement or, where the clear is a request's last statement, let the commit report success while silently discarding the request's earlier writes. A fourth, added by #511, has the same shape one level out: the shared rate-limit store is fail-open too. Any SQLAlchemyError, timeout or cross-loop RuntimeError yields "no verdict" and the process-local verdict stands, and the breaker then skips the store entirely for 60 seconds. So while a database fault persists, the shared buckets silently revert to per-worker limits — limit x worker count in the worst case — and a log line is the only evidence. Failing closed would 429 the whole site on a database blip, and unlike the lockout above there is still a real cap in force the whole time, which is why this trade is easier than the third. Row eviction under the counter table's cap is fail-open by the same logic and by construction: forgetting a counter can only allow a request, never falsely deny one.

Client-IP attribution (#509): X-Forwarded-For is honoured only when the socket peer is inside a configured trusted-proxy CIDR (DIOGENES_TRUSTED_PROXY_CIDRS, the ALB's subnets in production), and the client is then the right-most hop in the chain that parses and is not itself trusted. An untrusted peer, an absent or all-trusted chain, and an unparseable hop all fall back to the socket peer, so a forged header can neither mint a fresh bucket nor burn someone else's. Empty configuration means trust nothing. GET /api/v1/health is exempt from the API-query bucket so a misconfigured trusted set degrades limits rather than failing the load balancer's health probe.


Audit Capabilities

  • Full log export in JSON or CSV with combinable filters (fingerprint, document hash, date range).
  • Hash chain verification via POST /api/v1/log/verify.
  • Signed tree heads for Merkle inclusion proofs.
  • OpenTimestamps proofs for Bitcoin-anchored timestamps.
  • Audit trail queries with fingerprint and document scope.

Application audit logging (#512)

Finding H-6. Controls AU-2, AU-3, AU-12.

Before this, there was no audit-event pipeline in the application. The only request-level trace was uvicorn's access line ('%(levelprefix)s %(client_addr)s - "%(request_line)s" %(status_code)s') — no timestamp, no subject, no correlation ID, no outcome beyond the status code — so an exported CloudWatch line could not answer AU-3's what/when/where/who. The transparency log is deliberately not the destination: it is a public, append-only record of business events, and putting authentication telemetry there would publish every user's failed-login history and source IP permanently.

Records are one compact JSON object per line on the diogenes.audit logger (stdout), carrying timestamp, event, subject, source_ip, outcome, request_id, plus method, path, factor, reason, target, status and trace_id.

Four layers.

Layer Where What it knows
0 — request context server/request_context.py, middleware.RequestContextMiddleware correlation ID, method, path, lazily-resolved source IP
1 — credential/challenge adjudication server/auth.py, server/pin_enforcement.py, api/signing_pin.py, api/webauthn.py "this proof was rejected", with a factor
2 — route outcomes the API routers "a complete authentication succeeded"; "revocation of key X was denied"
3 — backstop RequestContextMiddleware's response hook any 401/403 nobody accounted for

Seven credential adjudicators, not two. AuthService.verify_challenge (signature), AuthService.check_password (argon2), verify_pin_binding at sign time and at rotation time, and three WebAuthn assertion sites — credential use, login, and PIN-management auth (Path B). The seventh was found by tests/test_audit_coverage_gate.py rule 4, not by reading, which is what that rule is for.

Terminality is not "has emitted". terminal_recorded means "a record has been written that is the account of this response". Only audit_log.denied — which builds the HTTPException the route raises — and an explicit emit(..., terminal=True) set it. emit never does, because deps.require_operator_jwt is a FastAPI dependency that resolves before the route body (so a success emit marking the request audited would suppress every later 403 on the operator routes), and because POST /api/v1/keys/revoke reaches Layer 1 through _verify_auth_proof (so a Layer-1 emit marking it audited would lose target — which key the revocation was denied for).

The backstop covers 401 and 403 only. 429 is excluded because an unmetered flood of audit lines would itself be the DoS, and #511 already meters every route so the 429 is the bounded refusal. 400 is excluded because it would sweep in ordinary Pydantic validation noise; the WebAuthn assertion failures that answer 400, and the challenge-state failures that answer 404/409/410, carry explicit Layer-1 records instead.

Two structural guarantees against credential leakage. emit has no **kwargs, so there is no channel through which a password, PIN, signature, token or nonce can reach a record; and reason/factor are closed StrEnums, so an exception string carrying input_value= '<secret>' — the H-5 shape — cannot be pasted into an audit line. subject and target carry only fingerprints, user UUIDs or pseudonyms, all already public, capped at 256 characters and JSON-escaped so a newline in a chosen pseudonym cannot forge a second line (AU-9).

There is deliberately no runtime kill switch: a disable knob on an AU-2 control is a finding, not a feature. The lever for volume is destination and retention.

Two records for one bad-signature login is intended. "This credential proof was rejected" and "this login attempt failed" are different facts at different layers, and AU-2 wants both.

Hand-off to #513 (H-7). This change delivers AU-2/AU-3/AU-12 capture to stdout. A retained, integrity-protected destination — object-locked S3, CloudTrail with log-file validation, AU-6 alerting, AU-11 retention — is #513's. trace_id (from X-Amzn-Trace-Id, honoured only from a trusted socket peer) is the join key and is inert until then.

L-18 (#523) is made visible, not fixed. api/federation.py still falls back to an anonymous UserContext when the operator key provider raises; it now emits operator.action / dev_unauthenticated so the fallback is auditable. Removing the fallback would change response codes on three federation routes and belongs to #523.