Security Architecture¶
This document describes the security architecture of Diogenes for security professionals evaluating the system's threat model, cryptographic design, and trust boundaries.
Design Principles¶
-
Subscriber private keys never leave the client. All subscriber signing is performed client-side: keys are generated in the browser with
crypto.subtle.generateKey(static/js/keys_register.js) and the server only ever receives public keys and signatures. That client-side generation is the SC-01 control.PrivateKeyRejectionMiddlewareis defence in depth, not the SC-01 control. It matches PEM markers in POST/PUT/PATCH request bodies only, and is trivially evaded by base64 or any other encoding. It catches an accidental paste; it stops nothing deliberate.One carve-out, by design: the log operator's own signing key is server-held. It is loaded from
operator.pemviaDIOGENES_OPERATOR_KEY_PATH(AWS Secrets Managerdiogenes/production/operator-keyin production) throughdeps.get_operator_key_provider()— seesrc/diogenes/server/api/endorsements.py. SC-01 scopes subscriber keys; the operator key is outside that scope and is protected by secret storage and IAM instead. -
Content never reaches the server. Documents are hashed locally. Only the SHA-256 digest is transmitted for verification. This preserves document confidentiality.
-
Append-only log. The transparency log is append-only and hash-chained. Once an entry is committed, it cannot be altered without breaking the chain. Chain integrity can be verified by any party.
-
No central authority. Trust is established through a decentralized web of endorsements. There is no certificate authority whose compromise would undermine the entire system.
-
Defense in depth. Multiple security layers operate independently: cryptographic verification, key status checking, trust assessment, rate limiting, password protection, and temporal anchoring.
Cryptographic Primitives¶
| Primitive | Usage | Notes |
|---|---|---|
| Ed25519 | Key pairs, attestation signatures | Recommended algorithm. Edwards-curve Digital Signature Algorithm. |
| ECDSA P-256 | Key pairs, browser signing | Compatible with Web Crypto API. SHA-256 hash. |
| RSA-2048 | Key pairs (legacy) | Supported for backward compatibility. |
| SHA-256 | Content hashing, fingerprints, log chain | Used throughout for hash computation. |
| Argon2id | Password hashing | Key passwords hashed with Argon2id before storage. |
| HMAC-SHA256 | JWT signing | JWTs signed with operator key for session management. |
Trust Boundaries¶
graph TD
subgraph Client["Client (Browser / CLI)"]
KG["Key Generation
(private)"]
SG["Signing
(private)"]
CH["Content Hashing
(document stays local)"]
end
subgraph Server["Server (API)"]
SV["Signature Verification"]
KR["Key Registry
(public keys)"]
TL["Transparency Log
(append-only)"]
RL["Rate Limiting"]
SH["Security Headers
(CSP, HSTS)"]
PKR["Private Key Rejection"]
end
subgraph DB["Database (PostgreSQL)"]
DATA["Keys, Log Entries,
Endorsements, Auth Challenges"]
end
subgraph BTC["Bitcoin (via OpenTimestamps)"]
TA["Temporal anchoring
of log entries"]
end
Client -- "HTTPS
(public keys, signatures, hashes only)" --> Server
Server --> DB
DB --> BTC
Authentication Model¶
Challenge-Response¶
Mutating operations (key succession, revocation, endorsement operations) require proof of key possession via a challenge-response protocol:
- The server issues a random nonce bound to a specific fingerprint.
- The client signs the nonce with their private key.
- The server verifies the signature against the registered public key.
Challenges expire (configurable TTL) and are single-use. Replay attacks are prevented by challenge consumption.
JWT Sessions¶
For browser-based portal sessions, a challenge-response exchange produces a JWT with:
- 30-minute TTL (configurable)
- 8-hour maximum session duration
- Fingerprint binding (JWT is valid only for the authenticated key)
- Optional recovery-only flag (limits operations available to recovered keys)
Password Protection¶
Keys can optionally register a password (Argon2id-hashed). When set, the password is required as a second factor for attestation signing and other sensitive operations.
Credential Binding¶
A session's identity is resolved purely by the credential row matching the JWT fingerprint. Binding a fingerprint to an account is therefore the identity grant itself, not a bookkeeping step — whoever can write that row can log in as that account.
Two consequences follow, and both are enforced rather than assumed:
- There is one binding path.
POST /api/v1/credentials/enrollrequires a signed challenge, an active non-revoked issuing credential, a verified chain signature, and hardware backing on both sides. The self-service/users/{user_id}/credentialsroutes that did the same job with no checks at all were deleted in #502 (C-2) rather than authenticated, because a hardened path and an unhardened path doing the same job is the shape that produced this finding in the first place. - Binding requires proof of the key, not just of the account.
CredentialService.enroll_credentialtakes a required proof of possession: either a signature by the key being bound, over canonical bytes naming both the account and the fingerprint, or an assertion that a WebAuthn ceremony has just proven possession for that same fingerprint. Deleting the routes alone would not have closed the finding — the next caller of the service would reintroduce it by omission.
The proof binds one (account, key) pair, so a proof captured for one account
does not verify for another and a proof made with one key does not verify for a
different one. It carries no nonce, deliberately: replaying it re-asserts the
identical binding, which is idempotent under the unique constraint on
credentials.key_fingerprint.
Known limit. Registration-time proof of possession is a separate gap:
POST /api/v1/keys/register does not prove the registrant holds the key it
registers (M-16, tracked in #518). Until that lands, the enrollment proof
above shows control of the registered public key, which is as strong as
registration made it.
Threat Model¶
Threats Addressed¶
| Threat | Mitigation |
|---|---|
| Server compromise | Private keys never reach the server. Existing signatures remain valid. The append-only log detects tampering via hash chain verification. |
| Log tampering | Hash chain integrity verification. Temporal anchoring to Bitcoin provides external proof of log state. Merkle tree heads enable efficient audit. |
| Key theft | Password protection adds a second factor. Key revocation is immediate and logged. Endorser revocation alerts notify the trust network. |
| Sybil attacks | Endorsement capacity scaling, activation delays, privilege thresholds, and over-capacity discounting prevent trust network gaming. |
| Replay attacks | Challenge-response nonces are single-use and time-limited. JWTs have short TTLs. |
| Content leakage | Documents are never uploaded. Only hashes are transmitted. |
| Rate-based abuse | Per-IP rate limiting with resource-specific thresholds (key registrations, endorsements, attestations, general queries). |
Threats Not Addressed (Out of Scope)¶
| Threat | Notes |
|---|---|
| Client-side key theft | Diogenes does not manage private key storage. Key holders are responsible for securing their private keys. |
| Network-level attacks | HTTPS is enforced in production (HTTPS redirect middleware), but network security is the deployment operator's responsibility. |
| Quantum computing | Current algorithms (Ed25519, ECDSA, RSA) are not post-quantum resistant. Algorithm agility in the protocol allows future migration. |
Security Headers¶
The SecurityHeadersMiddleware sets the following headers on all responses:
| Header | Value | Purpose |
|---|---|---|
X-Content-Type-Options |
nosniff |
Prevent MIME type sniffing |
X-Frame-Options |
DENY |
Prevent clickjacking |
Content-Security-Policy |
see below | Prevent XSS and injection |
Strict-Transport-Security |
max-age=63072000; includeSubDomains (HTTPS only) |
Enforce HTTPS |
X-XSS-Protection is not set. The header is retired in every current
browser and its legacy auditor was itself an injection vector; the row that
used to claim it here was simply wrong.
Content-Security-Policy¶
The application policy — every route that touches session JWTs, key material or user content — is emitted verbatim as:
default-src 'self';
script-src 'self' 'nonce-<per-request>' 'wasm-unsafe-eval';
style-src 'self' 'unsafe-inline';
font-src 'self';
img-src 'self' data:;
connect-src 'self' https://api.iconify.design https://api.simplesvg.com https://api.unisvg.com;
base-uri 'self';
form-action 'self';
frame-ancestors 'none';
object-src 'none'
The nonce is minted per request by SecurityHeadersMiddleware and rendered
into every inline <script> through a Jinja context processor. Because a
nonce is present, browsers ignore 'unsafe-inline' in script-src entirely —
so appending it back is not a rollback lever; reverting the policy commit
is. 'wasm-unsafe-eval' is retained for the vendored hash-wasm argon2id
derivation used by the signing PIN (#256); it permits WebAssembly compilation,
not JavaScript eval.
Since inline handler attributes cannot carry a nonce, no template contains an
on*= attribute: behaviour binds through data-on-* attributes and the
delegated dispatcher in static/js/actions.js. That is enforced structurally
by tests/test_template_csp_gate.py, which also forbids an inline <script>
in any components/ template — a component may be returned standalone as an
htmx fragment, on a different request with a different nonce.
Documented deviation. Three prefixes carry narrower policies that still
allow script-src 'unsafe-inline' and a CDN origin:
| Prefix | Why |
|---|---|
/api/docs, /api/redoc |
FastAPI generates the page, including an inline init script and cdn.jsdelivr.net asset URLs. Noncing it requires replacing docs_url/redoc_url with custom routes. |
/docs |
MkDocs build output. Its inline scripts change on every docs edit, the tree is not in the repo, and the Material theme pulls unpkg.com (mermaid), Google Fonts, and api.github.com for the repository facts in its header. |
All three are same-origin with the application, so the mitigation genuinely
does not extend to them. Their content is generated from repo-controlled
sources — our OpenAPI schema and our Markdown — and reflects no request input,
so there is no known injection path into them. Prefix matching is
exact-or-slash-prefixed, so a route such as /docsomething receives the strict
application policy. Closing the deviation means vendoring swagger-ui, redoc and
mermaid; the constants in middleware.py name that as their deletion
condition.
Headers are defence in depth, not the encoding control. nosniff does nothing
for a response that already declares text/html.
Output encoding¶
HTML is written by exactly one thing: the Jinja environment, which autoescapes
every .html template. Response bodies constructed in Python must therefore be
string literals — a body built from an f-string, a variable, or %/format
bypasses the template engine and with it the only encoding control.
That rule is enforced by tests/test_html_response_literal_gate.py (#506 /
H-5), which also asserts per template that autoescape is genuinely on and that
no template reaches for |safe, Markup(, or {% autoescape false %}.
Exception text reaches the user only through
diogenes.server.web._errors.safe_error_detail, which names the field and the
reason but never echoes the submitted value — pydantic's default str(exc)
repeats the caller's input verbatim in input_value=.
Rate Limiting¶
Rate limits are enforced by one pure-ASGI middleware that runs ahead of
routing, authentication and every dependency. Which routes are metered, and
against which budget, is data — server/rate_limit_policy.BUCKETS is the
single source of truth, and a test walks app.openapi() to fail the build if
any published route resolves to no bucket.
Coverage is a default with named exceptions, not a list. Before #511 the
middleware returned early for everything outside /api/ and dispatched
inside /api/ from a hard-coded if/elif over five paths, so 52 of the 57
API write routes and all 53 web routes were unmetered (finding M-10).
Shared marks a bucket whose count is authoritative across workers, held
in rate_limit_counters rather than in process memory. Before #511 nothing
was: docker-entrypoint.sh runs uvicorn --workers 4, so a configured limit
of N was really 4N per task and every deploy reset it (finding M-7).
| Bucket | Default limit | Shared | Purpose |
|---|---|---|---|
API queries (api-query) |
100/minute per IP | no | GET/HEAD/OPTIONS tier default under /api/ |
API writes (api-write) |
120/minute per IP | no | tier default for every other API method — the 52 routes nobody enumerated |
Key registrations (key-reg) |
10/hour per IP | yes | Limit Sybil key creation |
Endorsement offers (endorsement-offer) |
20/day per IP | yes | Limit trust network gaming |
Attestation events (attestation) |
50/hour per IP | yes | Prevent log flooding (/documents/sign shares it, so the alias is not a bypass) |
Log head anchoring (log-anchor) |
12/hour per IP | yes | Bound anonymous flood and OTS calendar amplification |
Account recovery (users-recover) |
10/hour per IP | yes | Bound guessing against an unauthenticated account-reset primitive |
Verification (verify) |
30/minute per IP | no | The seven unauthenticated verification POSTs plus the two portal forms: pure CPU (DSSE checks), no database |
Claim reveal (claim-reveal) |
30/hour per IP | yes | Bound the guessing rate against the claim commitment (does not close M-15 — the commitment is still unsalted) |
Web pages (web-page) |
240/minute per IP | no | GET/HEAD/OPTIONS tier default outside /api/ |
Web writes (web-write) |
60/minute per IP | no | tier default for every other portal method |
PIN proof checks (pin-check) |
5 per 15 min per fingerprint | no — already durable | Bound signing-PIN guessing; a pre-filter in front of the DB-backed signing_pins lockout |
| Authentication failures | 5 per 15 min per fingerprint | durable | Survives restart and is shared across workers (auth_lockouts) |
| Recovery failures | 5 per 15 min per account | durable | Survives restart and is shared across workers (auth_lockouts) |
Exempt, deliberately: GET /api/v1/health (a 429 there fails the ALB target
group's matcher = "200" and cycles the ECS task), /static/* and /docs/*
(StaticFiles mounts — no database, no application CPU, and one docs page
pulls dozens of assets). First match wins and buckets never stack, so a
request is charged to exactly one budget and Retry-After stays meaningful;
the consequence is that a client's total ceiling is the sum of the buckets it
touches.
The middleware runs ahead of the auth dependencies, so the log-anchor and account-recovery buckets also cap unauthenticated floods against those routes.
Shared limiter state (#511)¶
Shared buckets consult a Postgres sliding-window counter after the local
sliding-window log allows. The local tier keeps an already-refused client off
the database and can never produce a false 429, because its limit equals the
global limit. A denied request does not increment the shared counter, so a
flood cannot inflate it past the limit. Retry-After in the shared tier is
computed from the sliding-window estimate rather than the window edge.
Because this middleware is outermost, a store call acquires a pooled
connection before the handler acquires its own. Three guards bound that: an
in-flight cap of two per worker (at most 2 of 10 pooled connections), a 250 ms
asyncio.wait_for around acquire and execute, and a breaker that opens for
60 seconds after five consecutive errors, timeouts or sheds.
The volumetric buckets (api-query, api-write, verify, web-page,
web-write) are deliberately not shared: making them globally exact costs
a database statement per request, which is itself the amplification a DoS
control exists to prevent. Their boundary stays the task's concurrency and the
ALB. This is the recorded residual of M-7 and the hand-off to #431, where a
Redis store behind the same RateLimitStore protocol would make per-request
coordination cheap enough to extend sharing.
DIOGENES_RATE_LIMIT_SHARED_STORE=none restores exact pre-#511 per-process
limiting. Every limit is likewise an environment variable on the existing ECS
task definition, so retuning or disabling a bucket needs no image build and no
terraform apply.
Request body cap (#511)¶
PrivateKeyRejectionMiddleware buffers the whole body of every POST, PUT and
PATCH ahead of routing and authentication, and did so with no size check
(finding M-9) — an unauthenticated memory-exhaustion primitive against a
1 vCPU / 2048 MiB task shared by four workers. It now refuses an over-cap
Content-Length without a single receive(), and aborts a streamed body the
moment the running total crosses DIOGENES_MAX_REQUEST_BODY_BYTES (default
2 MiB, derived from the 1 MiB maximum payload plus DSSE envelope headroom)
rather than draining it. The 413 carries Connection: close; the residual is
that a client that already streamed megabytes may see a reset before it reads
the response, and draining politely would be the denial of service.
Durable authentication and recovery lockout (#510)¶
The last two rows are not middleware buckets. They are rows in
auth_lockouts, keyed (scope, subject) — auth on a key fingerprint,
recovery on a user id — and they mirror the counter columns signing_pins
already carries. Persisting them in the database rather than in process
memory is the point: an in-memory counter resets on every deploy and each
worker keeps its own, so five workers mean five budgets. A locked subject is
answered 429 with Retry-After, and the window self-expires — no operator
unlock exists or is needed.
The brake sits inside the two AuthService primitives, not on the routes.
The argon2 password is the genuinely brute-forceable credential (it is
human-chosen), and it is adjudicated at five places: POST /auth/login,
POST /auth/password, POST /api/v1/attestations, and — through
AuthService.require_auth — the key and endorsement routes. Guarding the
primitive covers all five and any sixth added later. The counter is cleared
by the routes instead, at the point a complete authentication succeeds: a
password brute-forcer supplies a valid signature on every attempt, so a
counter reset on signature success would oscillate and never arm.
POST /api/v1/auth/challenge is deliberately not guarded. It is the
shared proof-minting primitive for every challenge-gated route and for the
recovery-eligibility path, and fingerprints are public (transparency log, key
registry), so guarding it would let six requests deny a victim every
authenticated operation and their recovery path for 15 minutes, renewably,
for any key in the system. It would also buy nothing: verify_challenge
marks a challenge used with a flush that the request transaction rolls back
when the handler raises, so a single challenge already funds unlimited
guesses. The brake belongs on the guess.
Four trade-offs are accepted knowingly. Lockout is itself a denial of service: anyone who knows a fingerprint can lock it out of login, password change, signing, key operations and endorsements for 15 minutes. That is bounded by the self-expiring window, by the challenge route staying open so the victim's recovery path is never denied, and by the counter arming only on genuine credential guesses — never on a missing, expired, used or wrong-purpose challenge, nor on a "password required" refusal. And recovery answers 429 where an unknown pseudonym answers 401; that is not a meaningful enumeration oracle, because pseudonyms are published in the transparency log and the public key registry. Failures are recorded only for accounts that exist, so a flood of random pseudonyms cannot mint unbounded rows — the per-IP recovery bucket bounds that instead.
The third is fail-open on a database error. All three halves of the brake
degrade to silence rather than to a refusal: the check returns "not locked",
and the failure count and the post-success clear are logged and swallowed. So
for as long as a database error persists, the durable brake on login, password
change, every route behind AuthService.require_auth, and account recovery is
inert, and a log line is the only evidence. The alternative — failing closed —
is a total authentication and recovery outage on the same fault, which is why
the trade is taken; docker-entrypoint.sh runs alembic upgrade head before
uvicorn starts, so a task never serves traffic against a schema missing
auth_lockouts, and the per-IP middleware buckets in the table above keep
running independently of the database, so an unauthenticated flood is still
bounded while the durable half is dark. The degradation is deliberately
narrow: only a SQLAlchemyError is swallowed, so a programming error in the
control fails loudly in development and CI instead of going quietly dark in
production. Both DB-touching halves that run on the caller's session — the
check's read and the clear's delete — do so inside a SAVEPOINT, because a
failed statement aborts the enclosing Postgres transaction: without it a
swallowed lockout error would either resurface as an unrelated 500 on the
request's next statement or, where the clear is a request's last statement,
let the commit report success while silently discarding the request's earlier
writes.
A fourth, added by #511, has the same shape one level out: the shared
rate-limit store is fail-open too. Any SQLAlchemyError, timeout or
cross-loop RuntimeError yields "no verdict" and the process-local verdict
stands, and the breaker then skips the store entirely for 60 seconds. So while
a database fault persists, the shared buckets silently revert to per-worker
limits — limit x worker count in the worst case — and a log line is the only
evidence. Failing closed would 429 the whole site on a database blip, and
unlike the lockout above there is still a real cap in force the whole time,
which is why this trade is easier than the third. Row eviction under the
counter table's cap is fail-open by the same logic and by construction:
forgetting a counter can only allow a request, never falsely deny one.
Client-IP attribution (#509): X-Forwarded-For is honoured only when the
socket peer is inside a configured trusted-proxy CIDR
(DIOGENES_TRUSTED_PROXY_CIDRS, the ALB's subnets in production), and the
client is then the right-most hop in the chain that parses and is not itself
trusted. An untrusted peer, an absent or all-trusted chain, and an
unparseable hop all fall back to the socket peer, so a forged header can
neither mint a fresh bucket nor burn someone else's. Empty configuration
means trust nothing. GET /api/v1/health is exempt from the API-query
bucket so a misconfigured trusted set degrades limits rather than failing
the load balancer's health probe.
Audit Capabilities¶
- Full log export in JSON or CSV with combinable filters (fingerprint, document hash, date range).
- Hash chain verification via
POST /api/v1/log/verify. - Signed tree heads for Merkle inclusion proofs.
- OpenTimestamps proofs for Bitcoin-anchored timestamps.
- Audit trail queries with fingerprint and document scope.
Application audit logging (#512)¶
Finding H-6. Controls AU-2, AU-3, AU-12.
Before this, there was no audit-event pipeline in the application. The
only request-level trace was uvicorn's access line
('%(levelprefix)s %(client_addr)s - "%(request_line)s" %(status_code)s')
— no timestamp, no subject, no correlation ID, no outcome beyond the
status code — so an exported CloudWatch line could not answer AU-3's
what/when/where/who. The transparency log is deliberately not the
destination: it is a public, append-only record of business events, and
putting authentication telemetry there would publish every user's
failed-login history and source IP permanently.
Records are one compact JSON object per line on the diogenes.audit
logger (stdout), carrying timestamp, event, subject, source_ip,
outcome, request_id, plus method, path, factor, reason,
target, status and trace_id.
Four layers.
| Layer | Where | What it knows |
|---|---|---|
| 0 — request context | server/request_context.py, middleware.RequestContextMiddleware |
correlation ID, method, path, lazily-resolved source IP |
| 1 — credential/challenge adjudication | server/auth.py, server/pin_enforcement.py, api/signing_pin.py, api/webauthn.py |
"this proof was rejected", with a factor |
| 2 — route outcomes | the API routers | "a complete authentication succeeded"; "revocation of key X was denied" |
| 3 — backstop | RequestContextMiddleware's response hook |
any 401/403 nobody accounted for |
Seven credential adjudicators, not two. AuthService.verify_challenge
(signature), AuthService.check_password (argon2), verify_pin_binding
at sign time and at rotation time, and three WebAuthn assertion sites —
credential use, login, and PIN-management auth (Path B). The seventh was
found by tests/test_audit_coverage_gate.py rule 4, not by reading, which
is what that rule is for.
Terminality is not "has emitted". terminal_recorded means "a record
has been written that is the account of this response". Only
audit_log.denied — which builds the HTTPException the route raises —
and an explicit emit(..., terminal=True) set it. emit never does,
because deps.require_operator_jwt is a FastAPI dependency that resolves
before the route body (so a success emit marking the request audited
would suppress every later 403 on the operator routes), and because
POST /api/v1/keys/revoke reaches Layer 1 through _verify_auth_proof
(so a Layer-1 emit marking it audited would lose target — which key
the revocation was denied for).
The backstop covers 401 and 403 only. 429 is excluded because an unmetered flood of audit lines would itself be the DoS, and #511 already meters every route so the 429 is the bounded refusal. 400 is excluded because it would sweep in ordinary Pydantic validation noise; the WebAuthn assertion failures that answer 400, and the challenge-state failures that answer 404/409/410, carry explicit Layer-1 records instead.
Two structural guarantees against credential leakage. emit has no
**kwargs, so there is no channel through which a password, PIN,
signature, token or nonce can reach a record; and reason/factor are
closed StrEnums, so an exception string carrying input_value=
'<secret>' — the H-5 shape — cannot be pasted into an audit line.
subject and target carry only fingerprints, user UUIDs or pseudonyms,
all already public, capped at 256 characters and JSON-escaped so a newline
in a chosen pseudonym cannot forge a second line (AU-9).
There is deliberately no runtime kill switch: a disable knob on an AU-2 control is a finding, not a feature. The lever for volume is destination and retention.
Two records for one bad-signature login is intended. "This credential proof was rejected" and "this login attempt failed" are different facts at different layers, and AU-2 wants both.
Hand-off to #513 (H-7). This change delivers AU-2/AU-3/AU-12 capture
to stdout. A retained, integrity-protected destination — object-locked
S3, CloudTrail with log-file validation, AU-6 alerting, AU-11 retention —
is #513's. trace_id (from X-Amzn-Trace-Id, honoured only from a
trusted socket peer) is the join key and is inert until then.
L-18 (#523) is made visible, not fixed. api/federation.py still
falls back to an anonymous UserContext when the operator key provider
raises; it now emits operator.action / dev_unauthenticated so the
fallback is auditable. Removing the fallback would change response codes
on three federation routes and belongs to #523.