Skip to content

Operational Runbook ​

This runbook describes the tools and signals available to operate a BYOS instance. It covers the admin dashboard, logs, Slack notifications, uptime monitoring, and a monitoring checklist.

Admin dashboard ​

The admin dashboard is a web interface that reads data from the BYOS PostgreSQL database. It runs on port 53001 (prod-local). It has four pages.

Overview ​

The Overview page shows the proposal funnel for a selected time range (24 hours, 7 days, or 30 days):

  • Proposals received
  • Proposals sent to auction
  • Proposals won
  • Proposals settled
  • Proposals discarded
  • Rejection breakdown by reason
  • Penalty counts and total penalized amount

Use the Overview page to get a high-level picture of service health and activity.

Subsolvers ​

The Subsolvers page shows one row per connected sub-solver with:

  • Proposal counts (received, settled, reverted, rejected, penalized)
  • Win rate (settled / (settled + reverted))
  • Total penalties: count and amount
  • Buffer balance
  • On-chain escrow balance

Use the Subsolvers page to identify a sub-solver with an abnormal revert rate or growing penalty count.

Proposals ​

The Proposals page shows a paginated list of proposals. You can filter by sub-solver address and by status.

Possible statuses: submitted, active, rejected, simFailed, executing, settled, settleFailed, penalized, cancelled, expired.

Each row links to the proposal detail page. The detail page shows:

  • Full proposal fields (tokens, amounts, sub-solver, order UID)
  • Settlement transaction with a link to the block explorer
  • Penalty transaction (if applicable)
  • Tenderly replay link for reverted or simulation-failed proposals
  • Audit trail: a chronological list of all events with timestamps and raw payload data

Use the Proposals page to find and inspect a specific proposal. Use the Tenderly link on the detail page to debug a reverted settlement.

System ​

The System page shows operational metrics at the time of page load:

  • Pending penalties count: penalty transactions that have not yet been submitted on-chain

Use the System page to detect whether the penalty worker is falling behind. A growing pending penalties count indicates a problem — see the monitoring checklist.

Logging ​

The service uses structured logging via pino.

Log level ​

Set LOG_LEVEL to control verbosity. Accepted values: trace, debug, info, warn, error, fatal. Default: info.

Set JSON_LOGS=true to emit JSON-formatted logs. This is recommended when shipping logs to a cloud aggregator.

A change to either variable requires a service restart:

bash
docker compose -f docker-compose.prod-local.yml restart byos

Key log lines ​

Use these log lines to debug specific workers.

Validation worker

MessageLevelWhat it means
"validation tick"infoNormal tick. Fields: total, enqueued, expired.
"proposal rejected"infoProposal failed escrow check or simulation. Field: reason.
"proposal sim failed"warnSimulation returned an unexpected error.
"failed to snapshot live proposals"errorThe validation worker could not read from the database.

Penalty worker

MessageLevelWhat it means
"revert debit landed"infoPenalty transaction confirmed on-chain. Fields: id, subSolver, amount, tx.
"settlement cost lookup failed"warnCould not calculate the cost of a settlement. The worker will retry.
"keeps failing; giving up"errorA penalty operation exhausted all retries. The debit was not submitted.

Balance refresh worker

MessageLevelWhat it means
"balance refresh set is over half its cap"warnThe number of tracked sub-solvers is large. RPC call volume may be high. Fields: active, cap, requests.
"balance refresh batch failed"warnA multicall batch to fetch escrow balances failed.

Retention worker

MessageLevelWhat it means
"retention sweep dropped proposals"infoNormal cleanup. Field: deleted.
"retention sweep failed"errorThe retention worker could not delete expired proposals.

Audit worker

MessageLevelWhat it means
"failed to enqueue slack notification"warnAn audit event was saved but the Slack job could not be enqueued.

Slack notifications ​

Configure SLACK_TOKEN and SLACK_CHANNEL to receive notifications. Both variables must be set together. If one is missing, no notifications are sent.

Failed notification jobs retry up to 5 times with exponential backoff. If a notification does not arrive, check the logs for "failed to enqueue slack notification".

Events ​

EventTriggerMessage content
Service startedProcess startupChain ID
New sub-solver connectedFirst proposal from a sub-solver addressSub-solver address
Proposal settledSettlement transaction confirmedSub-solver, order UID, transaction hash
Settlement revertedSettlement transaction reverted on-chainSub-solver, order UID, transaction hash, Tenderly link
Sub-solver penalizedEscrow debited for a revertSub-solver, order UID, amount, penalty transaction hash
Sub-solver non-settlement debitedEscrow debited for abandoning a won auctionSub-solver, order UID, amount, penalty transaction hash
BYOS buffer debitedInternal buffer clearedSub-solver, amount, entries cleared, transaction hash

Redis persistence and the "new sub-solver" notification ​

The service tracks known sub-solver addresses in a Redis set. When a sub-solver sends its first proposal, the set does not contain its address, so a "new sub-solver connected" notification fires.

docker-compose.prod-local.yml enables AOF persistence (--appendonly yes) by default. If you are running Redis outside that compose file, ensure --appendonly yes is set — without it, every Redis restart will re-fire the "new sub-solver connected" notification for every known sub-solver.

Uptime monitoring ​

Use an external uptime tool (for example, BetterStack) to poll the health endpoint:

GET http://<host>:59585/healthz

A 200 OK response confirms the HTTP server is running.

Note. The /healthz endpoint is a liveness probe only. It does not check database connectivity, Redis, or chain connectivity. Use Slack alerts and logs as health signals for those dependencies.

Monitoring checklist ​

Check these items regularly to detect operational issues before they affect sub-solvers.

What to checkWhere to lookNormal stateAction if abnormal
Pending penalties countAdmin dashboard → System pageNear zeroCheck operator wallet balance; check logs for "settlement cost lookup failed" or "keeps failing; giving up"
Operator wallet balanceBlock explorer for the ESCROW_OPERATOR addressSufficient to cover gas for several penalty transactionsSend native tokens to the operator address
Slack alerts activeSlack channelPeriodic settled and settleFailed alerts during auction activitySilence here means the service is not winning auctions or not receiving /notify callbacks from the driver
Error-level log linesService logsNoneInvestigate the specific worker named in the log

This specification is normative. Where an implementation disagrees with it, the implementation is wrong.