Repo:github.com/pragnya-works/Edward · every claim below is read from the code (file paths in each section).
What this is: the complete path from "user sends a message" to "preview URL streams back", component by component, in the order the code actually runs.
Prod only — createForceHttpsMiddleware() reads x-forwarded-proto; anything not https is 308-redirected to https://host + originalUrl.
2. Admission control, step by step
Two layers, in this order: rate limit (Redis, cheap, per-scope) then admission (Postgres, authoritative, concurrency).
2a. Rate limits (middleware/rateLimit.ts)
Every limiter is built by createRateLimiterForScope(scope), which looks the policy up in RATE_LIMIT_POLICY_BY_SCOPE (@edward/shared/constants). If the policy is missing it throws at import time ("Rebuild @edward/shared…") — a guard against dist/source drift.
Store is rate-limit-redis on the same Redis connection, keyed rl:<policyPrefix>:, keyed by userId when authenticated (fallback: ipKeyGenerator).
On breach: set RateLimit-Scope, log security telemetry, return 429 with the policy's message.
The daily chat limiter is special — it is applied after the burst limiter and is a success-only quota: a separate daily-success counter (chatDailySuccess.service.ts) decides what actually burns the allowance, and are hand-set from a snapshot.
2b. Window check (getRunAdmissionWindow)
count() all rows in run whose status is in ACTIVE_RUN_STATUSES (queued, running) → activeRunDepth.
userRunLimit = min(MAX_ACTIVE_RUNS_PER_USER=2, AGENT_RUN_WORKER_CONCURRENCY=max(configured,2)) — the per-user cap can never exceed the worker's own concurrency.
overloaded = activeRunDepth >= MAX_AGENT_QUEUE_DEPTH (200). If overloaded → the request is rejected here with a stream 429 "under high load" and no run row is created.
2c. The admitted run (createRunWithUserLimit, packages/auth/lib/run.ts)
This is the step-by-step core you asked about:
Open one Postgres transaction.
Lock globally:pg_advisory_xact_lock(hashtext('run_admission_global')) — a single writer for the global count.
Lock per user:pg_advisory_xact_lock(hashtext(userId)) — this is why two messages fired at the same instant by the same user can't both pass.
(Transaction-scoped advisory locks release automatically on commit/rollback, so there is no lock leak if the request dies mid-way.)
Count active runs globally; if >= maxActiveRunsGlobal (200) → return {run: null, rejectedBy: "global_limit"}.
Count active runs for this user; if >= maxActiveRunsPerUser (2) → "user_limit".
Count active runs for this chat; if >= maxActiveRunsPerChat (1) → "chat_limit".
States a run walks through (RunState): INIT → LLM_STREAM → TOOL_EXEC → APPLY → NEXT_TURN → COMPLETE | FAILED | CANCELLED.
3. The queue, step by step (services/queue/*)
Two BullMQ queues on Redis (lib/queue.binding.ts): chat-processing-queue (BUILD + BACKUP) and agent-run-queue (AGENT_RUN).
Every payload is Zod-validated before it is added (BuildJobPayloadSchema.parse / AgentRunJobPayloadSchema.parse) — an invalid job can't enter the queue even from internal code.
Deterministic job ids = natural dedupe. Agent run: agent-run-<runId>. Build: build-<sandboxId>-<buildId|chatId:messageId> (identifier sanitized to [a-zA-Z0-9:_-]). Before adding, the code calls getJob(jobId); if it exists, it returns the job id instead of duplicating work.
4. The worker, step by step (apps/api/queue.worker.ts)
Boot: dotenv, then new Worker(BUILD_QUEUE_NAME, handler, { connection, concurrency: BUILD_WORKER_CONCURRENCY }) and the same for AGENT_RUN_QUEUE_NAME with AGENT_RUN_WORKER_CONCURRENCY.
Agent-run dispatch: the handler parses with JobPayloadSchema.parse(job.data); anything that is not AGENT_RUN on that queue is logged and thrown (no silent mis-execution).
Build dispatch:BuildQueueJobPayloadSchema.safeParse then a switch on BUILD / BACKUP; unknown types throw.
4a. Build job handler, step by step (workerJobHandlers.service.ts)
Resolve (or create) the build row; if it's already terminal (success/failed) → just re-publish its status and exit (idempotent duplicate handling).
Claim via conditional UPDATE:UPDATE build SET status='building' WHERE id=? AND status='queued' — only the winner gets a row back. Losers either see it terminal or "already claimed" and exit. This is the concurrency guard.
Publish building to edward:build-status:<chatId> (with retry).
Run buildAndUploadUnified(sandboxId) under withTimeout(WORKER_BUILD_JOB_TIMEOUT_MS).
On success: updateBuild(success, previewUrl, buildDuration), publish , then enqueue a job (non-fatal if that fails).
5. The agent-run worker: orchestration, step by step (services/runs/agent-run-worker/processor.ts)
getRunById(runId); if missing → throw; if already terminal → return (no double terminal transitions).
Terminal-event guard: if a session_complete meta event already exists for the run, map its termination reason to a status and finish without re-running the session.
Create the worker's AbortController and subscribe to the run's cancel channeledward:run-cancel:<runId>. A message on that channel aborts the stream — this is how the UI's Stop reaches a live worker in another process.
Start a status watchdog every RUN_TERMINAL_STATUS_POLL_INTERVAL_MS (2 s); if the run row became terminal elsewhere (another worker, a reaper), abort locally.
Parse run.metadata (parseAgentRunMetadata). Failure → mark run FAILED with the parse error.
Load the user's BYOK key and decrypt it (AES-256-GCM). Missing key / decrypt failure / unknown provider → FAILED with a user-facing message.
5a. What happens inside the session (services/chat/session/*)
Framework resolution — frameworkResolution.ts uses the workflow's intent: nextjs / vite-react / vanilla, or recovers the framework from an existing sandbox/Redis.
Prompt assembly — composePrompt() with a profile (DEFAULT/STRICT), modular system-prompt sections plus compacted skill packs.
Token budget — computeTokenUsage() per provider (tiktoken for OpenAI, provider counters for Gemini/Anthropic, reserved output tokens, vision overhead).
Turn loop — agentLoop.runner.ts / : max turns, max tool calls per turn (12) and per run (24) from ; each turn streams from with up to 2 stream attempts per turn.
6. The custom streaming parser, step by step (lib/llm/parser.ts, parser.shared.ts, parser.content.ts, parser.textSandbox.ts)
The model is instructed to emit tag-delimited blocks. The parser is a state machine over a rolling buffer, not a regex pass:
State starts TEXT. Buffer is capped at MAX_BUFFER_SIZE (10 KB) — on overflow the oldest bytes are dropped.
Each process(chunk) appends the chunk and then loops handleState() until the buffer stops shrinking or MAX_ITERATIONS (1000) is hit; hitting the cap emits a fatalparser_iterations_exceeded ERROR event and resets to TEXT (infinite-loop guard).
Tags come from ASSISTANT_STREAM_TAGS (): thinking start/end, start/end, file start/end (with dynamic path), install start/end, command, web search, url scrape, response, done.
7. Handling each event: writes, installs, commands (services/chat/session/events/*)
SANDBOX_START → if the workflow has no sandbox, ensureSandbox() provisions/reuses one and marks sandboxTagDetected (nothing else may touch the sandbox before this tag).
FILE_START → prepareSandboxFile(): acquire the flush lock, delete any Redis buffer for that path, mkdir -p the parent inside the container, truncate -s 0 the target. The in-memory generatedFiles map gets the path.
FILE_CONTENT → handleFileContent() strips a leading ``` fence on the first chunk, then (section 8). Content is also appended to in memory for post-gen validation.
8. Buffering the generated code and writing it into the container, step by step
This is the part most people miss: the model's bytes never go straight into Docker. They are debounced in Redis and flushed in batches.
Buffer keys (sandbox/write/shared.ts): edward:buffer:<sandboxId>:<path> (the content, appended) and edward:buffer:files:<sandboxId> (a set of dirty paths). TTL on both = SANDBOX_TTL (15 min).
writeSandboxFile()
Resolve sandbox state from Redis; normalize the path (reject absolute, .., empty) — a second path jail independent of the command one.
isProtectedFile(path, framework) — template files (package.json, tsconfig.json, , , , , …) are , so the model can't break the scaffold.
9. The Docker container: invocation, template auto-selection, isolation, step by step
9a. When it is invoked
The sandbox is provisioned lazily, on the first SANDBOX_START/file event of a run (ensureSandbox) or explicitly by the workflow's install/build phases. Nothing is pre-created per message.
The framework comes from the intent analyzer (an LLM classification, section 5a) or from an explicit preference detected in the user's words (frameworkPreference.ts), and is cached in Redis per chat (edward:chat:framework:<chatId>, 7-day TTL) so follow-ups reuse it.
9c. Provisioning, step by step (lifecycle/provisioning.ts)
waitForProvisioning(chatId) polls for an existing active sandbox (or a live edward:locking:provision:<chatId> lock) up to PROVISIONING_TIMEOUT_MS (30 s).
acquireDistributedLock('edward:locking:provision:<chatId>', 60 s) — SET NX PX + Lua CAS release, so exactly one provisioning runs per chat across all API/worker processes. Failing the lock → jittered retry (200–500 ms), up to 10 attempts, then "lock acquisition timeout".
Double-check getActiveSandbox(chatId) under the lock (someone may have finished while we waited) and return it if present.
nanoid(12) sandbox id; lifecycle state → PROVISIONING (Redis lifecycle record).
9d. The network window (how packages install in an isolated container)
The container is created network-less and verified so.
INSTALL_PACKAGES (steps/executeInstallPhase.ts) calls connectToNetwork(containerId), then mergeAndInstallDependencies(...) (pnpm), then always calls disconnectContainerFromNetwork(...) — success or failure, and even if the install threw (the error is captured first, disconnect runs after).
buildAndUploadUnified does the same: it connects at the start, may install pnpm globally / run pnpm install --frozen-lockfile=false if node_modules is missing, and disconnects on return path including the catch block.
10. Build: validation, unified build, upload, step by step
Package resolution first (RESOLVE_PACKAGES, dependency.resolver.ts + registry/package.registry.ts): normalize specs, block a blocked-package list, validate each package against the npm registry, cache results in Redis, and mark each valid/invalid. The valid set becomes resolvedPackagesbefore anything is installed.
Install (9d) with retries: INSTALL_PACKAGES = 3 retries, 120 s timeout (workflow/config.ts).
processBuildPipeline (in-session, right after the loop):
11. "If an error occurs during build, how does it get fixed" — the five fix layers, step by step
The repair strategy is layered, cheapest first. Nothing here is "hope the model notices".
Deterministic in-place autofixes (no model involved) — postgenAutofix.ts:
rewrites broken zustand default imports into { create as X } form,
forces <link rel="canonical"> hrefs in index.html to absolute https URLs (a href="/" makes Vite's build-html read a directory and crash),
the sanitize step (section 7.4) strips CDATA/fences/HTML entities from every written file.
Static post-gen validators — planning/validators/postgenValidator.*: framework, imports, logic, SEO and constants. Each violation carries a severity; recoverable ones are streamed to the user as [Validation] …, and ones are collected into a blocking report.
12. Saving: Postgres (truth) vs Redis/Upstash (live), step by step
12a. Postgres — durable truth
run — every run: id, chatId, userId, userMessageId, assistantMessageId, status, state, currentTurn, nextEventSeq, loopStopReason, , , , timestamps. State is updated at each phase (, , , , …) and on checkpoint.
12b. Redis (Upstash-compatible, single shared instance) — live state
The API and worker talk to one Redis through lib/redis.ts (ioredis, REDIS_URL host+port parsed in app.config.ts, maxRetriesPerRequest: null). Every key family:
Purpose
Key
TTL
Run event fan-out (pub/sub)
edward:run-events:<runId>
—
Build status fan-out
edward:build-status:<chatId>
—
Run cancel signal
edward:run-cancel:<runId>
—
File buffers
edward:buffer:<sandboxId>:<path>
15 min
Dirty file set
edward:buffer:files:<sandboxId>
12c. S3 + the backup path
Build output is uploaded per user/chat prefix (buildS3Key(userId, chatId, 'preview/…')), and the sandbox source is archived (source_backup.tar.gz, gz + tar) on a BACKUP job and on cleanup, with a Redis flag and an S3 HeadObject check so restore can find it later.
13. Streaming back to the user, step by step
Two paths run in parallel on purpose: Postgres for correctness, Redis for liveness.
Producer side (worker): for every parsed event, persistRunEvent(runId, event, publisher) does appendRunEvent (Postgres, gets seq) and then publisher.publish("edward:run-events:<runId>", {id, runId, seq, eventType, event}) (Redis). The DB is the truth; Redis is the live fan-out. If the subscriber is down, nothing is lost — the DB replay still has it.
Consumer side — the same request that POSTed the message stays open and unifiedSendMessage ends by calling streamRunEventsFromPersistence({req, res, runId}). So /chat/message itself is the SSE stream; the client never has to poll.
Two deployment modes (EDWARD_DEPLOYMENT_TYPE, auto-resolved to subdomain when complete Cloudflare config exists, else path).
14a. What the API does after a successful upload
Files land in S3 under <sanitizedUserId>/<sanitizedChatId>/preview/… with sanitized key components and content types mapped by extension.
CloudFront cache invalidation — invalidatePreviewCache(userId, chatId) creates an invalidation for /<userId>/<chatId>/preview/* via @aws-sdk/client-cloudfront (non-fatal on failure).
Subdomain assignment (previewRouting/subdomain.ts): if the chat has no customSubdomain yet, one is generated deterministically from sha256(userId:chatId) → a unique-names-generator adjective-animal name plus a 5-char base36 suffix (attempt 2+ mixes in the attempt number). It's claimed with a conditional and retried up to 5 times on a unique-constraint collision. User-chosen subdomains are validated first: length 3–63, , and a reserved-word list.
14b. What the Cloudflare Worker does on every request (scripts/worker.js, wrangler.toml)
Read env.CLOUDFRONT_URL; if unset → 500 "Worker misconfigured".
Split the hostname; if there is no subdomain or it's in the RESERVED set (www, api, admin, app, mail, dashboard, ftp, dev, smtp, staging, preview, static, assets, cdn, media, files, storage) → plain 200 "Welcome to Edwardd". This is what keeps the apex/marketing and reserved hosts out of the preview path.
Only GET/HEAD are allowed → anything else is 405.
Rate limit per subdomain via the Durable Object: idFromName(subdomain) → POST /check. The DO keeps a per-minute window (Math.floor(Date.now()/60000)) and a count in its own storage (loaded in blockConcurrencyWhile), with a cap; over the cap → 429 + seconds to the next minute boundary. If the DO binding is missing or errors, it fails open.
End-to-end subdomain flow: user prompt → build → S3 preview/ → CloudFront invalidation → KV subdomain → userId/chatId → a request to https://<subdomain>.edwardd.app/<route> hits the Worker → DO rate check → KV lookup → CloudFront fetch of <prefix>/preview/index.html or the asset → hardened response with the iframe CSP.
15. One-paragraph recap of the whole chain
POST /chat/message → Helmet/CORS/1 MB → Zod → Redis rate limits → getRunAdmissionWindow → createRunWithUserLimit under two Postgres advisory locks (global, user) and three caps (200 global / 2 per user / 1 per chat) → run row queued/INIT → BullMQ agent-run-queue (job id agent-run-<runId>) → worker picks it up, subscribes to the cancel channel, decrypts the BYOK key, reconstructs history → runStreamSession decides the framework, assembles the prompt, computes the token budget and streams model output through the custom tag parser → typed events drive sandbox provisioning (a locked, network-isolated, non-root node:22-slim container from the framework template), file writes (Redis-buffered, debounce → cat >> into the container), sanitize, installs (only inside a briefly opened network window) and commands (allow-list + path jail) → autofixes, validates and enqueues a build → worker claims the build row with a conditional UPDATE, runs syntax/types/build validation in the container, fixes via strict re-generation / RECOVER, diagnoses failures into a structured report → artifacts upload to S3, CloudFront invalidates, the subdomain is claimed in Postgres and mapped in Cloudflare KV → every event is written to Postgres ( with a monotonic seq) published to Redis → the still-open SSE stream replays Postgres from while live Redis events are buffered and merged by seq, with backpressure and heartbeats → the browser renders the stream and can resume mid-build → the Cloudflare Worker resolves through the Durable-Object limiter and KV to CloudFront, enforcing path safety, caching and the iframe CSP.
16. Gaps / things worth knowing
Path vs subdomain preview are both live paths in the code; which one you get depends on complete Cloudflare config (EDWARD_DEPLOYMENT_TYPE wins, otherwise auto).
The DEPLOY phase is a formality in stepRunner.ts — it just checks that state.context.previewUrl exists. The real publishing happens in the build orchestrator's upload/registration step.
Preview URLs are only registered for chats with userIdandchatId on the sandbox; otherwise upload/cleanup is skipped with a warning.
Several non-fatal steps are deliberately swallowed (CloudFront invalidation, SPA fallback upload, stale-preview cleanup, post-build backup enqueue) — the build still succeeds.
Redis is the single point of coordination: locks, buffers, rate limits, pub/sub and BullMQ all share it. If it's unavailable, admission still works (Postgres) but queueing, streaming fan-out and file flushing do not.
The container pins User: "node" on create on every exec, and fails closed — a container that can't be isolated is destroyed rather than used.
CORS — explicit allow-list (dev allows all), credentials on, Last-Event-ID explicitly allowed and the RateLimit-* headers explicitly exposed. This is what makes the SSE resume work cross-origin.
Body caps — express.json({ limit: "1mb" }) and urlencoded 1 MB, so a huge payload dies before any service code.
Telemetry — securityTelemetryMiddleware issues or adopts x-request-id and logs 401/403/429/5xx as http_anomaly. That id is the traceId carried into the run metadata.
Route chain for a message —
POST /chat/message → authMiddleware → chatRateLimiter → dailyChatRateLimiter → validateRequest(UnifiedSendMessageRequestSchema) → unifiedSendMessage.
Zod itself — validateRequest(schema) calls schema.parse(req) on the whole request (body/params/query), then writes the parsed values back onto req (this is why unknown keys are stripped and types are coerced). A ZodError returns 400 with { error, details: [{path, message}], timestamp }. A malformed payload never reaches a service.
RateLimit-Limit / -Remaining / -Reset
Insert the run row: id = nanoid(24), status: "queued", state: "INIT", plus chatId/userId/userMessageId/assistantMessageId/metadata.
Commit and return the run.
Back in the orchestrator: on rejection the freshly saved user message is deleted (cleanupUnqueuedUserMessage) and the user gets a specific message — global load, "this chat already has an active run", or "too many active runs, limit=N".
On success: a second META SSE frame with the runId is written to the client, then enqueueAdmittedRun(runId).
If enqueue itself fails → the run is immediately marked failed with the error message (never left dangling in queued).
existing
Retry/backoff policy per job type: build = 3 attempts, exponential from 2 s; backup = 2 attempts, fixed 1 s; agent run = 2 attempts, fixed 1.5 s.
History caps:removeOnComplete count 100 (build) / 200 (agent run); removeOnFail 50 / 100 — the queue doesn't grow forever.
On enqueue failure the throw is caught and the caller (build pipeline / admission) writes a terminal FAILED state plus a published status event, so the UI never waits on a job that will never run.
Two background loops are started inside the worker process (not a separate service):
processScheduledFlushes() every 250 ms — drains the deferred sandbox file-flush markers.
reapStaleRuns() every 60 s — kills runs stuck in queued older than 10 min or running older than 45 min.
Lifecycle hooks are registered for logs/failures; SIGINT/SIGTERM/uncaughtException/unhandledRejection all funnel into createGracefulShutdown with a 12 s budget, which closes both workers, quits the Redis publish client and clears both intervals.
A separate publish client (createRedisClient) is used for pub/sub so a blocked subscriber can never stall job execution.
success
backup
On failure: build a structured error report from the container (createErrorReportIfPossible → services/diagnostics/*), persist failed + errorReport, publish failed, rethrow so BullMQ records the failure.
Publish helper (workerPolicies.ts): publish with timeout + N retries to edward:build-status:<chatId>; give up with a warning rather than hanging.
assertModelMatchesProvider(keyProvider, metadata.model) — a GPT key can't be used with a Gemini model id.
Follow-up runs reconstruct history in the worker (buildConversationMessages(chatId)) rather than trusting the metadata snapshot; on failure it falls back to the snapshot.
Build the progress object (turn, first-token latency, stop reason, termination reason, error) and markRunRunningIfAdmissible(runId) — the queued → running transition is itself conditional, so a cancelled run can't be resurrected.
Capture the stream instead of a real socket:createRunEventCaptureResponse(onEvent) returns a fake Express Response that buffers bytes, splits them on \n\n into SSE frames, keeps the data: payloads, JSON-parses them into StreamEvents and feeds them, serially, to the persistence callback. It decodes with StringDecoder so multi-byte UTF-8 split across chunks can't corrupt the stream.
Call runStreamSession(...) — the whole session runtime (section 6) — with the capture response, the abort signal, the decrypted key, history and onCheckpoint (which persists currentTurn + resumeCheckpoint back onto the run).
capturedRes.flushPending() drains the serialized persist queue (and rethrows the first persistence error), then finalizeSuccessfulRun.
Finalize success: re-read the run; if it already went terminal, do nothing. Otherwise write status/state/currentTurn/loopStopReason/terminationReason/errorMessage/completedAt plus metadata (runDurationMs, firstTokenLatencyMs, resumeCheckpoint: null) and log the run_completion metric.
Finalize failure: classify the error (classifyAssistantError), drain pending events, then explicitly persist an ERROR event and a session_complete meta event with terminationReason: stream_failed, then mark the run FAILED. If that terminal write fails, the original error is rethrown.
finally: clear the watchdog, unsubscribe and quit the cancel subscriber.
agentLoop.stream.ts
loop/budgets.ts
streamResponse(...)
Every chunk goes through the parser (section 6) and each parsed event is handled (section 7) and written to the response, which the worker capture turns into persisted events.
Continuation/never-truncate policy — shared/continuation.ts with MAX_AGENT_CONTINUATION_PROMPT_CHARS (14 000) and MAX_NEVER_TRUNCATE_CHARS (100 000).
After the loop — buildPipeline.processBuildPipeline() runs the autofix → validation → build-enqueue sequence (sections 10–11).
parser.shared.ts
edward_sandbox
Attributes (e.g. the file path) are read with extractTagAttribute, which first normalizes escaped quotes — LLM output frequently emits \".
No-op closing tags like </edward_command>, </edward_web_search>, and any unexpected </edward_*> are dropped instead of being printed to the user; </edward_install> / </edward_sandbox> are preserved because they close real state.
Lookahead of LOOKAHEAD_LIMIT (256) bytes means a tag split across two network chunks is still recognized — no partial-tag leakage.
flush() on stream end emits the residual buffer as the correct type and closes any open state (FILE_END + SANDBOX_END, INSTALL_END, THINKING_END) so a truncated response still produces balanced events.
writeSandboxFile()
generatedFiles
FILE_END → sanitizeSandboxFile(): the API runs a small Node script inside the container to strip stray <![CDATA[…]]> / ``` fences and unescape HTML entities (<, &, …). This is a real deterministic repair pass, not a prompt hope.
SANDBOX_END → flushSandbox() — the buffered code lands in the container before anything is built.
COMMAND → tool gateway (events/tools/command.ts → sandbox/command.service.ts): a command allow-list (ls, find, grep, mv, cp, mkdir, rm, cat, pnpm, npm, git, pwd, date, echo, touch, head, tail, wc, tsc), a disallowed-pattern filter (rm -rf /, writes into /etc, chmod, chown), arg count ≤ 60, arg length ≤ 1024, total args ≤ 8 KB, control-char rejection, find -exec banned, and a path jail that normalizes every path-like argument and refuses anything outside /home/node/edward (and refuses rm on the workdir root). Execution is User: "node", cwd /home/node/edward, 45 s default timeout.
WEB_SEARCH / URL_SCRAPE → websearch service; outbound fetches go through an SSRF-aware safe fetch.
INSTALL_* → queued through an install task queue so package installs are serialized and awaited before the build phase.
Every event is written to the (captured) response, which the worker turns into a persisted run_eventand a Redis publish (section 12).
next.config.*
vite.config.*
tailwind.config.*
globals.css
refused
One Redis pipeline: APPEND content, SADD the dirty set, PEXPIRE both. Pipeline errors are collected and thrown as one error.
STRLEN the buffer: over MAX_WRITE_BUFFER (5 MB) → immediate flush; otherwise → debounced flush.
Debounce (flush.scheduler.ts): scheduleSandboxFlush(sandboxId, immediate) writes edward:flush:due:<sandboxId> = now (+WRITE_DEBOUNCE_MS 100 ms if not immediate). If immediate, it kicks flushSandbox() right away.
The worker tick runs processScheduledFlushes() every 250 ms: SCAN for edward:flush:due:* keys (skipping :lock), compare the due timestamp to now, then take a short distributed lock (SCHEDULER_LOCK_TTL 5 s) so only one process handles a sandbox at a time. The marker is deleted before flushing; on failure it's re-set for a retry.
flushSandbox() (the actual write into the container)
Take the flush lock edward:flush:lock:<sandboxId> (SET NX PX, TTL 30 s, Lua compare-and-delete release). waitForLock=true retries up to 20× 250 ms — this is what SANDBOX_END and the final build flush use.
Load sandbox state, get the container, ensureContainerRunning.
Loop (max MAX_FLUSH_FAILURES 10): read the dirty set; for each path atomically claim it with RENAME edward:buffer:… → …:processing and SREM from the set, so a concurrent writer re-dirties the path instead of losing bytes.
GET the processing key, then open a hijacked exec running sh -c "cat >> '<fullPath>'" with AttachStdin: true, write the buffered content to the stream and end it. Single-quotes in the path are shell-escaped ('"'"').
Inspect the exec exit code (non-zero is logged), then DEL the processing key in a finally — so nothing is double-appended.
If a claim failed, count a failure and retry after 500 ms; at 10 failures it throws "Flush failed after maximum retries".
Also on flush:prepareSandboxFile uses truncate -s 0 for a rewrite, clearBuffers/cleanupBufferKeys clean up on teardown, and cleanupSandbox flushes before backing up and destroying the container so no bytes are lost at expiry.
TEMPLATE_REGISTRY picks the image and the output dir: nextjs → out, vite-react → dist, vanilla → ., each …/nextjs-sandbox:latest etc. from DOCKER_REGISTRY_BASE. Unknown framework → getDefaultImage() (vanilla).
FRAMEWORK_CONTRACTS pins the dependency set and required scripts per framework (Next 16 / React 19 / Tailwind 4 / TypeScript 5.9 …) and can validate() a generated package.json.
If the framework is still unknown at build time, detectFrameworkFromPackageJson() reads the container's package.json and persists the detected framework on the sandbox (and Redis).
Labels: com.edward.sandbox=true, com.edward.user, com.edward.chat, com.edward.sandboxId (these labels are how a container is recovered after a Redis loss).
container.start(), then disconnectFromNetwork("bridge") and verifyNetworkIsolation() — inspect the container's interfaces; if anything is still attached, the container is destroyed and provisioning fails loudly.
No Privileged, no extra capabilities, non-root enforced twice (USER node in the image andUser: "node" on every create and every exec).
Optional restore from S3 backup if shouldRestore and a backup exists (Redis flag first, then S3 HeadObject).
ensureScaffoldGitignore() base64-writes a default .gitignore (skipped for vanilla, non-fatal on failure).
saveSandboxState() → Redis keys edward:sandbox:<id> and edward:chat:sandbox:<chatId> (both 15-min TTL, refreshed on use) + framework key.
Lifecycle → ACTIVE; release the lock. On any failure the lifecycle record goes FAILED and the lock is always released.
Reuse path:getActiveSandbox first checks a cached container-status key (10 s cache), then container.inspect(), refreshing the TTL on success and cleaning up dead containers (FAILED state, delete Redis state). If Redis state is gone entirely, it falls back to scanning Docker containers by label and rehydrating Redis from the labels.
Expiry:SANDBOX_TTL 15 min, refreshed on every use; cleanupExpiredSandboxContainers() runs every CLEANUP_INTERVAL_MS 60 s; cleanup flushes → backs up → destroys → deletes Redis state → clears buffers (state CLEANING_UP → TERMINATED).
every
So the only moment untrusted generated code can reach the network is the package-manager window; the sandbox is deaf the rest of the time.
If there are no generated files and no declared packages → skip the build entirely.
validateGeneratedOutput({ framework, files, declaredPackages, mode }) → post-gen validators (framework/style/SEO/imports/logic). Every violation is streamed to the user as a recoverable [Validation] … error.
Blocking error-severity violations → the run does NOT enqueue a build. A build row is created directly as FAILED with a structured errorReport, that report is published to edward:build-status:<chatId> and a build_status event is sent. (This is why some failures are instant and never enter the queue.)
flushSandbox(sandboxId, waitForLock=true) — final guarantee that all buffered code is in the container.
Otherwise create a queued build row, publish queued, and enqueueBuildJob(...). An enqueue failure is turned into a FAILED build row + published report, not a silent hang.
The build job (4a) calls buildAndUploadUnified(sandboxId):
Connect to network; verify pnpm exists (install it globally as root if not — the one intentional root exec).
Auto-detect the framework from package.json if still unknown.
If node_modules is missing (e.g. a restored sandbox) → pnpm install --frozen-lockfile=false with NEXT_TELEMETRY_DISABLED=1 CI=true NPM_CONFIG_ENGINE_STRICT=true.
mergeAndInstallDependencies(...) for the run's requested packages.
runValidationPipeline(containerId, sandboxId) — three real stages inside the container:
syntax: node --check on every .ts/.tsx/.js/.jsx under src.
types: pnpm tsc --noEmit piped through head -50 (60 s).
build: pnpm run build --if-present, with an EXIT_CODE:$? marker so the real exit code survives the pipe (120 s).
Any failure returns — and the retryPrompt is a generated instruction block naming the stage and up to 10 concrete errors.
runUnifiedBuild(...) produces the output dir; then upload: uploadBuildFilesToS3(...), SPA fallback 404.html uploaded for non-vanilla frameworks, stale preview keys cleaned (cleanupS3FolderExcept), invalidatePreviewCache() (CloudFront invalidation of /<userId>/<chatId>/preview/*, non-fatal), and finally the preview URL is resolved — subdomain mode calls registerPreviewSubdomain(...), path mode uses buildPathPreviewUrl(...).
Disconnect from the network on every path; return {success, buildDirectory, previewUrl, previewUploaded, error?}.
error-severity
Strict re-generation — runStreamSession.strictRetry.ts: if blocking violations ≥ STRICT_RETRY_MIN_VIOLATIONS (4) and a sandbox exists, the session re-runs the agent loop once with a STRICT prompt profile and a retry prompt built from the actual violations (buildPostgenRetryPrompt). It snapshots the current files/packages first so a failed retry can be rolled back to the previous generation.
Workflow RECOVER phase — workflow/config.ts + engine.ts: every phase (ANALYZE 2, RESOLVE 3, INSTALL 3, GENERATE 2, BUILD 3, DEPLOY 2) has a retry budget. A failed step that still has retries left moves the workflow to RECOVER (LLM-executed, 60 s); RECOVER can then jump back to the step after the last successful one — i.e. the engine re-enters the pipeline where it broke rather than restarting. Retries use exponential backoff capped at 10 s (retry.ts). If retries are exhausted or RECOVER itself fails → workflow FAILED.
Diagnosed error report for the user — queue.worker.helpers.ts → services/diagnostics/*: the raw build error is parsed into typed errors (missing_import, type_mismatch, syntax, config, runtime, resource, network, environment) and each gets probableCause / pinpointContext / preciseFix / nextStep — including special-cased diagnoses like "Vite crashed inside its own bundle → Node/Vite version mismatch, do not edit node_modules". This report is persisted on the build row and published in the build_status event, so the UI can show why it failed, not just that it did.
terminationReason
errorMessage
metadata
LLM_STREAM
TOOL_EXEC
APPLY
NEXT_TURN
run_event — every stream event, appended through appendRunEvent: inside one transaction it does UPDATE run SET nextEventSeq = nextEventSeq + 1 … RETURNING, then inserts run_event with id = <runId>:<seq> and the same seq. The nextEventSeq counter on run is what makes the sequence monotonic and unique per run without a race; replay reads getRunEventsAfter(runId, afterSeq, 500).
run_tool_call — every tool attempt: turn, toolName, idempotencyKey, input, output, status (started|succeeded|failed), errorMessage, durationMs, with a unique (runId, idempotencyKey) constraint and an upsert — so a retried tool call reuses its row instead of duplicating.
message / chat — the user message is saved before admission (and deleted if admission rejects); the chat row also carries customSubdomain.
build — build records with status, previewUrl, errorReport, buildDuration.
15 min
Flush scheduling
edward:flush:due:<sandboxId> (+ :lock)
15 min / 5 s
Flush mutex
edward:flush:lock:<sandboxId>
30 s
Provision mutex
edward:locking:provision:<chatId>
60 s
Sandbox state / chat index
edward:sandbox:<id>, edward:chat:sandbox:<chatId>
15 min
Cached framework
edward:chat:framework:<chatId>
7 days
Backup flag
edward:backup:exists:<chatId>
7 days
Container liveness cache
runtimeState store
10 s
Lifecycle state
runtimeLifecycle store
—
Rate limiters
rl:<scope>:<key>
per policy
BullMQ queues
chat-processing-queue, agent-run-queue
—
Cache-Control: no-cache
Connection: keep-alive
configureSSEBackpressure(res) installs a per-response writer with a queue and a drain handler, so a slow client doesn't block the worker's event pipeline.
Subscribe-before-replay (the important ordering): the subscriber attaches to edward:run-events:<runId>first, buffering any live events in a Map<seq, envelope> while replaying = true. Only then does it query Postgres getRunEventsAfter(runId, lastSeq, 500) in batches and emit them in seq order. Once the replay loop finds no more rows, it flushes the buffered live events, de-duplicating anything seq <= lastSeq and sorting by seq. Result: no gap and no duplicates even if events landed while the DB query was in flight.
Resume after a reload: the client sends Last-Event-ID (query param or header; the CORS list allows it). parseLastEventSeq accepts either a bare number or the <runId>:<seq> id form, clamps absurd values (>50 000 000 → 0). The server then starts from that seq — so a page refresh mid-build resumes exactly where it stopped, which is the whole reason events were persisted at all.
Frames:sendSSEEventWithId(res, eventId, event) writes id: <runId>:<seq>\ndata: <json>\n\n (version stamped STREAM_EVENT_VERSION). Keep-alive comments (: run-events-heartbeat) are sent during replay every 10 s and there's a 15 s heartbeat interval, so proxies don't cut an idle stream.
Termination: a meta event with phase: session_complete is the terminal marker. When it's emitted, terminalEventSeen = true and the stream is closed with sendSSEDone (data: [DONE]) and the connection ends. Client disconnect (req.on("close")) unsubscribes and ends quietly; the run itself keeps going in the background (the worker already owns it — the comment in the orchestrator says exactly this).
Backpressure/limits: replay is capped at 500 rows per batch; a dropped write marks the client slow and closes the stream rather than corrupting it; the whole response is a real Express response so Node's socket backpressure applies underneath.
Client: the web app consumes the SSE in apps/web/stores/chatStream/* (useStartStream, resumeRunStream, cursorPersistence), dispatches parsed events into chat state (chatStreamProcessor), and persists the last event id per chat/run so it can resume. Build status arrives on a separate channel (edward:build-status:<chatId> → build sync hooks) rather than inside the run stream.
UPDATE chat … WHERE id=? AND customSubdomain IS NULL
^[a-z0-9][a-z0-9-]{1,61}[a-z0-9]$
KV mapping write (kvClient.ts): PUT https://api.cloudflare.com/client/v4/accounts/<account>/storage/kv/namespaces/<ns>/values/<subdomain> with the storage prefix (<userId>/<chatId>) as the value — a 10 s AbortController timeout, errors thrown with status + body slice.
Availability check (checkSubdomainAvailability) verifies DB ownership and the current KV value (so a subdomain can't be stolen from another chat); deleting a chat deletes the KV entry, and only if the current KV value still matches that chat's prefix.
500 requests/minute
Retry-After
KV lookup:env.SUBDOMAIN_MAPPINGS.get(subdomain) → the storage prefix. No value → the styled 404 "App not found" page.
Path safety: reject NUL, backslashes, .. segments, //, and control chars, both raw and percent-decoded; a decode failure is also rejected.
Routing decision: a path with a file extension is an asset → /<prefix>/preview<pathname>; anything else is a route → /<prefix>/preview/index.html (this is the SPA-deep-link fallback).
Fetch from CloudFront (${CLOUDFRONT_URL}${cfPath}) with Accept-Encoding passthrough and its own User-Agent.
If a non-asset returns 403/404 → the same 404 page (so a deleted preview doesn't leak CloudFront's XML).
Headers on the way out: content type from upstream, Cache-Control = public, max-age=31536000, immutable for assets (or no-store on failure) and public, max-age=60, must-revalidate for HTML, X-Content-Type-Options: nosniff, Referrer-Policy: strict-origin-when-cross-origin, and content length when present.
CSP frame-ancestors rewrite: for text/html only, the Worker takes the upstream CSP and force-upserts frame-ancestors (default 'self' http://localhost:3000 https://edwardd.app https://www.edwardd.app, overridable by env.FRAME_ANCESTORS) — that's how the built app is allowed to be embedded in the Edward UI iframe while everything else stays locked down.
The body is streamed straight through (new Response(response.body, …)), so the Worker never buffers whole files.