BareGit
# Reliable HTTP API and message delivery

Status: proposed. Date: 2026-09-21.

## 1. Purpose and scope

Complete the existing Telegrammer API by making successful responses truthful,
invalid requests predictable, and background processing recoverable. This
proposal covers the current working tree, including its uncommitted API-key
implementation. It supersedes conflicting response behavior in `design.md`.

The existing `POST /send` and `POST /subscribe` endpoints remain. Sending stays
synchronous; callbacks retain the original Telegram Message JSON body. SQLite
continues to store API credentials. Use libmw for HTTP, URL parsing, and SQLite.

This design additionally proposes persistent subscriptions, a durable delivery
queue, and subscription removal. Those are deliberate extensions to the original
memory-only implementation: retries and restarts otherwise cannot provide a
useful delivery guarantee. They should ship as a separately reviewable phase.

Not included: arbitrary Telegram API proxying, media uploads, multiple bots in
one process, active-active servers, exactly-once delivery, or a management UI.
All API keys retain the existing authority to send to any chat the bot can reach.
Subscription ownership controls management, not per-chat access rights.

## 2. Findings and required outcomes

| Current behavior | Required outcome |
| --- | --- |
| `/send` returns 200 for parsed upstream error JSON | 200 requires upstream success |
| Upstream JSON exceptions escape polling | Typed failures and contained threads |
| `ok: false` polling responses loop immediately | Observable, bounded backoff |
| Callback HTTP errors are ignored | Classify status and retry eligible failures |
| Detached thread per callback | Fixed workers and durable pending work |
| Shared ordinary `bool` controls polling | Coordinated stop and joined threads |
| Callback URLs and numeric IDs are loosely converted | Validate before side effects |
| Duplicate subscriptions accumulate | Idempotent registration per owner |
| Keys use a PRNG seeded from one random-device result | 256 bits of secure randomness |
| Keys are stored and listed in plaintext | Digest storage and one-time disclosure |
| Packaged service has no writable database directory | Explicit service state directory |
| Shell smoke script prints output without assertions | Repeatable automated tests |

The prior review confirmed a successful build and local 401, 400, and subscribe
200 responses. It did not verify real Telegram delivery or callback recovery.
Its plaintext-key concern is a hardening proposal, not a violation of the old
PRD, which only required SQLite storage. SQLite sharing alone is also not a
confirmed race: libmw documents a serialized connection. Transaction ownership
still needs explicit handling as described below.

## 3. HTTP contract

### 3.1 Common request processing

Process requests in this order:

1. Match method and path. Unknown paths return 404; recognized paths with an
   unsupported method return 405 and an `Allow` header.
2. Authenticate before buffering the body using the HTTP server's pre-request
   hook. Reject missing, malformed, or unknown bearer credentials with 401 and
   `WWW-Authenticate: Bearer`. Accept the scheme case-insensitively; the token
   remains case-sensitive. Reject multiple Authorization headers.
3. Enforce a 64 KiB body limit, including chunked requests, returning 413.
4. Require `application/json`, allowing a charset parameter, otherwise 415.
5. Parse a JSON object and validate all fields before any outbound request or
   database write. Syntax, type, range, and field errors return 400.
6. Execute the operation and translate its typed result to an HTTP response.

Use JSON for all API errors:

```json
{
  "ok": false,
  "error": {
    "code": "INVALID_REQUEST",
    "message": "chat_id must be a nonzero signed integer",
    "field": "chat_id"
  }
}
```

`field` is optional. Stable uppercase codes are for programs; messages are for
people. Never return raw exception text, token-bearing URLs, SQL, or complete
upstream bodies. An authentication database failure is 503 `AUTH_UNAVAILABLE`,
not 401. Every response uses `Cache-Control: no-store`.

Configure server read/write timeouts and bounded request concurrency. Defaults:
5 seconds for request reads/writes, 8 HTTP workers, and 64 queued requests.
Verify the pinned httplib version's settings and rejection behavior; do not
assume a queue limit exists merely because the pool has a fixed size.

### 3.2 POST /send

Request example: `{"chat_id":123456789,"text":"Hello"}`. Alternatively use
`{"username":"some_user","text":"Hello"}`.

Require exactly one of `chat_id` and `username`. This intentionally replaces
the old behavior where `chat_id` silently won when both were supplied.

`chat_id` must be an integer in the signed 64-bit range and nonzero. Negative
IDs are allowed. Reject floats, booleans, numeric strings, and unsigned values
above `INT64_MAX`; do not let JSON numeric conversions truncate or wrap them.
`text` must be a string, nonempty, valid UTF-8, and at most 4096 Unicode code
points. This endpoint does not accept parse mode. Telegram remains the final
authority on text validity. Reject unknown request fields to expose typos.

`username` accepts one optional leading `@`, then ASCII letters, digits, and
underscores, with a 64-byte defensive limit. Normalize ASCII case for lookup.
Unknown mappings return 404 `USERNAME_NOT_FOUND`; update the old design's 400
documentation. Store mappings only from private-chat messages, so activity in
a group cannot redirect a message addressed to a person into that group.
When a known private chat changes or removes its username, remove the old
association. Persist mappings with their chat ID and observation time. They
are observations, not authoritative directory lookups; renames not yet seen
by the bot can leave stale data. Clients needing a stable destination use IDs.

A successful response preserves Telegram's successful JSON envelope:

```json
{"ok":true,"result":{"message_id":42,"chat":{"id":123456789}}}
```

The example abbreviates the actual Message object. Return the complete result.
Only an upstream 2xx response with boolean `ok: true` and an object `result`
qualifies. Validate this before responding 200.

| Failure | Local status and code |
| --- | --- |
| Local request invalid | 400 `INVALID_REQUEST` |
| Username absent | 404 `USERNAME_NOT_FOUND` |
| Telegram rejects destination or content (400/403) | 422 `TELEGRAM_REJECTED` |
| Telegram rate limit | 429 `TELEGRAM_RATE_LIMITED` |
| Telegram token invalid (401), other error, invalid envelope | 502 `UPSTREAM_ERROR` |
| DNS, connection, TLS failure | 502 `UPSTREAM_UNAVAILABLE` |
| Transfer timeout | 504 `UPSTREAM_TIMEOUT` |
| Unexpected internal failure | 500 `INTERNAL_ERROR` |

Recognize Telegram errors from both HTTP status and the JSON envelope. Include
validated positive `retry_after` seconds as `Retry-After` for rate limits.
Bound it to an integer representable by the scheduler. Never reinterpret a
Telegram 401 as a failure of the caller's gateway key.

Do not automatically retry `sendMessage`. A timeout may occur after Telegram
accepted the message, so another attempt can duplicate it. Document this
ambiguity for callers; durable outgoing sends are outside this proposal.

### 3.3 POST /subscribe

Require exactly `chat_id` and `callback_url`. Apply the same ID rules. Validate
URLs with `mw::URL`: absolute HTTP or HTTPS, nonempty host, valid port, maximum
2048 bytes, no user information, fragments, control characters, or whitespace.
Retain path and query because callbacks may require them. Never log the query.

Allow loopback and private addresses: local callbacks are a core requirement.
All key holders are therefore trusted to choose reachable callback targets.
Restrict initial and redirect protocols to HTTP/HTTPS and disable redirects.
This is not a public, untrusted webhook registration service. Deployments needing
destination restrictions can add an explicit address allowlist later.

Normalize scheme/host case and default ports using the parser, but do not
reorder query parameters or decode/re-encode paths. Use the resulting canonical
URL for uniqueness together with owner and chat ID.

Insert the subscription in a transaction. Repeat registration by the same key
for the same chat and URL returns the existing ID and creates no new delivery
stream. Both initial and repeated requests return 200:

```json
{"ok":true,"subscription_id":17}
```

Registration does not contact the callback or prove that the bot sees the chat.
Success means the subscription was committed. It starts receiving updates
processed after that commit, not historical deliveries.

Add `GET /subscriptions` to list only the caller's subscriptions, with IDs,
chat IDs, and URLs. Add `DELETE /subscriptions/{id}` to remove an owned
subscription and its pending deliveries, returning 204. Unknown or foreign
IDs return 404. A callback already in flight may finish after deletion; no
later retry should be scheduled. Both additions require authentication.

## 4. Components and ownership

Split `main.cpp` into focused components as implementation progresses:

| Files | Responsibility |
| --- | --- |
| `api_server.h/.cpp` | Routes, validation, HTTP result translation |
| `telegram_client.h/.cpp` | Telegram requests and response interpretation |
| `key_store.h/.cpp` | Credential creation, lookup, migration, revocation |
| `subscription_store.h/.cpp` | Subscription ownership and idempotency |
| `delivery_store.h/.cpp` | Transactional ingestion, cursor, delivery jobs |
| `poller.h/.cpp` | Long polling, retry policy, ingestion |
| `dispatcher.h/.cpp` | Fixed callback workers and completion handling |
| `application.h/.cpp` | Startup, stop coordination, ownership |
| `main.cpp` | CLI parsing and exit status |

Use `mw::E<T>` with concrete error types carrying category, optional upstream
status, retry delay, and a sanitized diagnostic. Do not use exceptions for
ordinary network or validation failures. Catch parsing exceptions at the
Telegram boundary and unforeseen exceptions at every thread entry point.

`Application` owns components through values or `unique_ptr`; workers borrow
references whose owners outlive their joined threads. Inject HTTP session
factories, a clock, and wait/backoff control for testing. Each concurrent
request gets its own session; do not share `mw::HTTPSession` across threads.

Use separate SQLite connections for worker transactions and request operations.
A connection belongs to one operation at a time. A serialized SQLite connection
does not prevent unrelated threads from accidentally joining the same explicit
transaction. Use short transactions and never hold one over network I/O.

Use `std::jthread`/stop tokens and interruptible condition-variable waits. A
fixed worker pool takes the place of detached callback threads.

## 5. Persistence and migrations

The current development database is not schema-versioned: its
`PRAGMA user_version` is `0`, and the existing layout is the legacy
`api_keys(name,key,created_at)` table. This implementation therefore creates
the new development schema directly, does not set or bump `user_version`, and
refuses the legacy table and other incompatible target-table layouts with a
clear error. Remove or replace that development database before starting this
build. A nonzero `user_version` is also rejected because no migration path is
implemented yet. Before production release, add an explicit `user_version` migration
sequence and a SQLite backup step; production migration is intentionally out
of scope for this development implementation. Restrict future backups to the
service user, since the old backup contains usable credentials.

Target schema (timestamps are integer Unix seconds):

```sql
CREATE TABLE api_keys (
    id INTEGER PRIMARY KEY,
    name TEXT NOT NULL UNIQUE,
    key_digest TEXT NOT NULL UNIQUE,
    created_at INTEGER NOT NULL
);
CREATE TABLE subscriptions (
    id INTEGER PRIMARY KEY,
    owner_id INTEGER NOT NULL REFERENCES api_keys(id) ON DELETE CASCADE,
    chat_id INTEGER NOT NULL,
    callback_url TEXT NOT NULL,
    created_at INTEGER NOT NULL,
    UNIQUE(owner_id, chat_id, callback_url)
);
CREATE TABLE poll_state (
    singleton INTEGER PRIMARY KEY CHECK(singleton = 1),
    bot_id INTEGER NOT NULL,
    next_offset INTEGER NOT NULL
);
CREATE TABLE usernames (
    username TEXT PRIMARY KEY,
    chat_id INTEGER NOT NULL UNIQUE,
    observed_at INTEGER NOT NULL
);
CREATE TABLE deliveries (
    id INTEGER PRIMARY KEY,
    subscription_id INTEGER NOT NULL
        REFERENCES subscriptions(id) ON DELETE CASCADE,
    update_id INTEGER NOT NULL,
    payload TEXT NOT NULL,
    state TEXT NOT NULL CHECK(state IN ('PENDING','IN_FLIGHT','DEAD')),
    attempt_count INTEGER NOT NULL DEFAULT 0,
    next_attempt_at INTEGER NOT NULL,
    created_at INTEGER NOT NULL,
    last_error TEXT,
    UNIQUE(subscription_id, update_id)
);
CREATE INDEX deliveries_due ON deliveries(state, next_attempt_at, id);
```

Enable foreign keys, WAL, `synchronous=NORMAL`, and a finite busy timeout on every
connection. On transaction failure roll back the whole operation. A failed
commit must not change the in-memory cursor. These boundaries rely on SQLite's
[transaction semantics](https://www.sqlite.org/lang_transaction.html) and
[atomic commit guarantees](https://www.sqlite.org/atomiccommit.html).

With WAL and `NORMAL`, committed state survives application crashes, but recent
commits may roll back after an OS crash or power loss. Database consistency is
preserved. Accept this durability tradeoff to avoid synchronizing the WAL on
every transaction; power-loss durability is not a requirement for this service.
See SQLite's [synchronous pragma](https://www.sqlite.org/pragma.html#pragma_synchronous).
Here, durable jobs means persistence across application crashes and restarts,
not guaranteed survival of machine failure. A lost ingestion commit after
Telegram has acknowledged its offset can lose callbacks; a lost completion
write can cause duplicate delivery. Recent credential and subscription changes
can also roll back under the same failure conditions.

Only one running daemon may use this state database. Open `${db}.lock` with
`O_CREAT|O_RDWR|O_CLOEXEC|O_NOFOLLOW`, mode `0600`, then acquire
`flock(LOCK_EX|LOCK_NB)` and retain the descriptor for the whole daemon
lifetime. Treat `EWOULDBLOCK`/`EAGAIN` as an actionable "another daemon is
running" error, and do not delete the lock file on shutdown. CLI credential and
delivery-management commands deliberately do not acquire this daemon lock;
SQLite's busy timeout and WAL handle their short operations. Validate the bot
identity through `getMe` before polling. Reject a database associated with a
different bot instead of applying its cursor and subscriptions to that bot.

Generate keys from 32 cryptographically secure random bytes and encode as
64 hexadecimal characters. Prefer an appropriate libmw crypto primitive if
available at the pinned revision; otherwise use OpenSSL `RAND_priv_bytes`,
checking its return value and failing closed. See the official
[random-byte API](https://docs.openssl.org/3.6/man3/RAND_bytes/).

Store SHA-256 of the exact bearer token, not the token. A fast digest is suitable
for newly generated high-entropy credentials; these are not human passwords.
Use a supported libmw hashing primitive or OpenSSL EVP. Print a new key only
once from `--add-key`; `--list-keys` prints names and creation times. Check bind,
query, and RNG errors explicitly. Reject empty names and names over 128 bytes.

The future production migration should hash existing plaintext keys exactly,
preserve their validity, and remove the old column transactionally. Existing
keys retain their old entropy and should be rotated; hashing does not
strengthen them. Explain that removing plaintext columns does not erase old
backups, WAL files, or storage remnants. Do not claim cryptographic erasure.

`--delete-key` deletes owned subscriptions and queued deliveries in the same
transaction. Already accepted sends and in-flight callbacks may finish. No
positive credential cache is needed initially, so later requests observe
revocation directly through the database.

## 6. Polling and ingestion algorithm

Use `getUpdates` with a 30-second long-poll timeout, a 5-second connection
timeout, a 40-second transfer timeout, and a 16 MiB response limit. Request
`allowed_updates: ["message"]` explicitly. Channel posts and edited messages
are outside this version's callback contract. Telegram only supplies updates
visible to the bot; do not promise messages hidden by its privacy configuration.

Telegram confirms updates through a later offset, so persistence must precede
that later request. See [getUpdates](https://core.telegram.org/bots/api#getupdates)
and [Update](https://core.telegram.org/bots/api#update).

For each successful batch:

1. Validate the envelope, result array, update IDs, and supported message
   structure before modifying state. Unknown extra fields are allowed.
2. Begin a database transaction. Read the current stored cursor and ignore
   updates older than it. Process the remaining updates in ascending ID order.
3. For each supported message, update private username observations and insert
   a job for every subscription visible in this transaction. Copy the message
   body as JSON without changing its fields. Uniqueness prevents duplicate jobs.
4. Advance `next_offset` to the last processed update ID plus one. Updates with
   valid IDs but unsupported types still advance the cursor without jobs.
5. Commit all jobs, mappings, and cursor together. Only then issue the next
   Telegram request using the new offset, and notify callback workers.

If a required field in a supported update is malformed, fail the batch, log a
sanitized diagnostic, and back off without advancing. This favors visibility
over silent loss; a persistent malformed update requires operator diagnosis.

On transport errors or upstream 5xx, use exponential backoff starting at one
second and capped at 60 seconds, with jitter. On rate limits wait at least the
validated server retry delay. A valid successful batch resets backoff. Invalid
credentials or polling conflicts stop polling, mark readiness false, and log
an actionable error; do not silently delete an existing Telegram webhook.

Set a default 100,000-job limit, counting pending, active, and dead jobs. Reject
an ingestion batch transaction if its new jobs would exceed capacity; leave the
cursor unchanged, mark readiness false, and retry after capacity is available.
Also limit subscriptions to 100 per key and 1,000 globally. Enforce limits
transactionally; an idempotent registration does not consume another slot.
Return 409 `SUBSCRIPTION_LIMIT` when a new registration exceeds its limit.

The queue avoids application-crash loss after commit. It cannot promise unlimited
retention while polling is stopped: Telegram itself has a finite update
retention window. Alert on sustained ingestion failure and queue saturation.

## 7. Callback delivery and recovery

Start four callback workers by default. Each worker claims one due pending job
in a short transaction, marks it `IN_FLIGHT`, and increments its attempt count.
Do not select work for a subscription with another active delivery. Prefer its
oldest non-dead job; a delayed retry blocks later messages for that subscription.
Other subscriptions can continue independently.

POST the stored Message JSON with `Content-Type: application/json`. Add
`X-Telegrammer-Delivery-Id` using a stable value derived from bot ID,
subscription ID, and update ID. Preserve that value across attempts so receivers
can deduplicate. Require a private database to prevent ID reuse through resets;
do not promise identity stability after intentionally replacing state.

Use a 3-second connection timeout, 10-second total timeout, and 64 KiB response
limit. Disable redirects. Do not log callback response bodies or URL queries.

| Result | Action |
| --- | --- |
| 2xx | Delete the completed job |
| Network error, timeout, 408, 429, or 5xx | Schedule retry |
| Other 4xx or 3xx | Mark `DEAD` with a sanitized reason |
| Retry budget exhausted | Mark `DEAD` |

Allow eight total attempts and at most 24 hours of retries. Backoff is
`min(300, 2^(attempt_count - 1))` seconds plus 0–25% jitter. For a valid
`Retry-After` delta or HTTP date, wait at least that long; if it exceeds the
remaining job lifetime, mark dead. Store due times in the database, but use
monotonic clocks for in-process waiting and recompute after clock changes.

After acquiring the daemon lock on restart, reset all `IN_FLIGHT` jobs to
`PENDING`; enforce budgets before another attempt. A crash after the recipient
accepts a message but before the completion write can cause a duplicate.
Delivery therefore uses bounded retries with possible duplicates, not an
unconditional at-least-once guarantee. A receiver should persist deduplication
state together with its application effect before returning 2xx.

Provide CLI `--list-failed-deliveries` (IDs, times, reasons),
`--retry-delivery <id>` (reset budget and state), and
`--delete-delivery <id>` for explicit operator recovery. Retain dead jobs for
seven days, then purge them with an aggregate warning and count. Automatic
purging is a documented retention policy; operators requiring longer retention
must configure it. No payloads appear in ordinary logs or listing output.

## 8. Startup, shutdown, and packaging

Validate configuration, database access, the development schema, and the
process lock before starting workers. `ApiServer::start()` registers HTTP
routes exactly once before its synchronous bind; `main()` does not call
`setup()` separately.

The pinned libmw revision exposes the underlying httplib bind/listen split, but
its convenience wrapper can still spin after a bind failure. `ApiServer` binds
synchronously with that underlying API, disables `SO_REUSEPORT`, and only then
starts the serving thread; an occupied port therefore returns a startup error
without an arbitrary sleep. Its serving thread is joined during shutdown.

On SIGINT/SIGTERM, use a dedicated POSIX signal-wait thread or a self-pipe.
Never invoke logging, database operations, or server methods directly inside
an asynchronous signal handler. Stop accepting requests, request worker stop,
wake retry waits, then join HTTP handlers, polling, and callback workers before
destroying their dependencies. Workers finish bounded active requests and claim
no new work. A 50-second shutdown budget covers the 40-second poll timeout;
verify that the HTTP server's stop/join behavior respects this budget.

Add authenticated `GET /health` returning component state and queue counts.
Return 503 when ingestion is suspended by a persistent failure or capacity;
transient retrying can report degraded state with 200. Never expose credentials,
callback URLs, or payloads. Define persistent failure as five consecutive
failures or five minutes without a successful poll; a successful poll clears it.

Default binding becomes `127.0.0.1`; preserve explicit `--host 0.0.0.0` for
remote deployments. Plain HTTP remains suitable for loopback; remote bearer
traffic requires TLS termination by the deployment.

Update the systemd unit with `StateDirectory=telegrammer`, `UMask=0077`, an
explicit `--db /var/lib/telegrammer/telegrammer.db`, and
`TimeoutStopSec=60`. Document how operators supply the token and provision keys
against that exact database as the service user. Declare SQLite as a runtime
package dependency. Pin libmw to the tested commit rather than tracking its
default branch, and pin any new dependency explicitly.

## 9. Compatibility and rollout

Publish these changes before rollout: errors become JSON; subscribe success
becomes JSON; both destination fields are rejected; malformed numeric values
are rejected; username lookup targets observed private chats; keys are no
longer listed; the default bind address changes; subscriptions become durable.
Existing valid bearer strings and successful send envelopes remain valid.

For this development transition, stop the old daemon, preserve any needed
backup, remove or replace the old plaintext-key database, and provision new
keys. There is no historical subscription or queue state to recover from the
old executable. Before production release, add the versioned migration and
backup procedure described in Section 5. Rollback requires the backup and the
old binary; do not run the old executable against migrated state. Reverting
loses new queue state, so drain or explicitly account for pending deliveries
before rollback.

## 10. Implementation sequence

1. Extract request validation and Telegram result interpretation behind injected
   HTTP clients. Establish tests before changing observable success behavior.
2. Implement HTTP status/error mapping, numeric and URL validation, explicit
   timeouts, contained polling failures, and synchronized shutdown. Fix and pin
   the libmw startup behavior. Deliver this as the core correctness phase.
3. Add secure key generation, digest lookup, revocation semantics, and service
   database configuration. Reject the old development schema explicitly; add
   production migration and rollback tests before release.
4. Add persistent owned subscriptions, list/delete endpoints, private username
   state, and bot identity binding. Document compatibility changes.
5. Add atomic cursor/job ingestion, capacity limits, callback workers, retries,
   crash recovery, operator commands, and readiness information.
6. Update README, original API specification, package dependencies, and the
   smoke script; complete end-to-end and failure testing before release.

Follow repository conventions: snake_case files/variables, CapCase types and
tests, camelCase functions, uppercase enum values, four-space indentation,
braces on new lines, and intention comments on public interfaces. Use `-j24`
for compilation. Keep storage, network, and protocol changes separately
reviewable; do not combine them into another monolithic `main.cpp`.

## 11. Verification and acceptance criteria

Register tests with CTest and run them without real Telegram credentials.
Use scripted fake HTTP clients for semantic errors and loopback HTTP servers
for transport behavior. Only tests can replace the Telegram endpoint with an
HTTP fixture; production always uses the configured bot's HTTPS API endpoint.
Use temporary databases and deterministic clocks; never touch the operator's
database or bot during automated tests.

| Area | Required cases |
| --- | --- |
| Authentication | Missing/duplicate/bad headers, valid/revoked key, DB failure |
| Validation | Invalid JSON, arrays/null, unknown fields, oversized/chunked body |
| IDs | Floats, booleans, strings, overflow, valid negative IDs |
| Send | 2xx success, `ok:false`, invalid result, HTML response, timeout, 429 |
| Usernames | Case/@ normalization, group isolation, rename/removal, restart |
| Subscriptions | Concurrent duplicate registration, ownership, limits, deletion |
| URL policy | Invalid URL, userinfo, non-HTTP scheme, redirect not followed |
| Ingestion | Crash before/after commit, replay, unsupported update, DB busy |
| Capacity | Batch rollback at limit, no cursor advance, recovery after drain |
| Callbacks | 2xx, 3xx, 4xx, 5xx, timeout, retry delay, budget, deduplication ID |
| Recovery | Crash after remote acceptance, active-job reset, failed-job CLI |
| Revocation | Pending jobs cancelled, accepted/in-flight behavior documented |
| Lifecycle | Occupied port, stop during poll/backoff/callback, thread joining |
| Migration | Legacy schema is rejected in dev; future migration preserves keys |
| Packaging | Service-user startup with explicit writable state directory |

Use barriers or fixture controls to crash the process at transaction and
delivery boundaries, rather than hoping timing reproduces failures. Verify
that jobs and cursors change atomically and that duplicate callbacks retain
their delivery ID. Run ThreadSanitizer for lifecycle and store concurrency,
and AddressSanitizer/UndefinedBehaviorSanitizer for parsing and recovery tests
where dependencies support them.

Build with `cmake --build build -j24`, then run
`ctest --test-dir build --output-on-failure`. CTest discovering zero tests is
a release failure. The shell smoke script must assert response status/body,
avoid printing credentials, use a temporary fixture instance, and fail on an
unexpected result. An optional manual real-bot check can verify sending and
private-message callbacks, but is not a substitute for deterministic tests.

Release requires: no false send success; no uncaught background exception;
bounded network and shutdown waits; committed jobs survive application restart;
retry and
revocation semantics match this contract; migration preserves existing valid
keys; and all registered tests pass.