A reliable writing product needs more architecture around the language model than inside the model call.
A reliable writing product needs more architecture around the language model than inside the model call. The model can produce paragraphs, but it cannot, by itself, establish who requested them, preserve progress through a deployment, enforce a credit balance, or explain why a publishing attempt failed.
That changes what you're building. A single API call can demonstrate generation. A product has to manage the work before, during, and after that call.
WriterzRoom's architecture reflects that wider responsibility. Its web application handles interaction and sessions. Its Python backend controls admission, credits, and a staged writing workflow. PostgreSQL stores content, progress, and checkpoints. Separate services handle specialized tools, while scheduled jobs recover work that an interrupted process cannot finish.
The central argument is straightforward: separating planning, drafting, editing, and formatting makes long-form generation easier to control because each stage has a specific responsibility and an inspectable result. The infrastructure around those stages makes the workflow usable as a service.
Part 2 follows that progression from the single-model starting point to the current layered system, then examines the operational changes that followed. Where release dates aren't established, the sequence follows architectural dependencies rather than claiming a dated development history.
A single-model prototype provides a useful baseline. You send a prompt, wait for a response, and display the text. For a developer testing whether a model can produce a coherent article, that loop answers an immediate question.
Long-form work quickly introduces questions the loop cannot answer. Which sources should the draft use? Can the user pause for review? What happens if generation finishes but publishing fails? How should the service respond when a user has already paid credits and the process disappears?
These questions shaped the responsibilities visible in WriterzRoom's current design. Generation is a workflow with persistent state, account permissions, external dependencies, and recovery rules. Each responsibility needs somewhere to live.
The repository contains two runtimes. frontend/ holds the Next.js application, using the App Router, React, TypeScript, and Tailwind CSS. langgraph_app/ holds the Python application, with FastAPI serving requests and LangGraph coordinating generation. The Python entry point is main.py.
Keeping both runtimes in one repository lets developers change the user interface and backend contract together. It doesn't make their responsibilities interchangeable. The frontend manages interaction and session authentication; the backend decides whether work is permitted and executes it.
That division also accommodates different development needs. An editor needs responsive controls, stable forms, and readable progress. A generation workflow needs structured state, source retrieval, model clients, database access, and error handling. TypeScript and Python serve those different concerns without requiring every component to use the same language.
You can read the architecture as a sequence of commitments. First, define the work. Then authorize it. Next, run it through explicit stages. Persist enough information to inspect and recover it. Finally, deploy the service without assuming that any individual process will remain alive.
The first architectural expansion is separating the work that a single prompt tries to perform simultaneously.
A long prompt might ask a model to research a topic, choose an angle, draft an article, edit its language, apply formatting, and prepare publication metadata. If the result is weak, the failure is difficult to locate. The draft may be poorly written because the plan was vague, because the evidence was unsuitable, or because the formatting instructions competed with the writing instructions.
WriterzRoom expresses these responsibilities as a LangGraph StateGraph over EnrichedContentState. The workflow moves through planning, research where applicable, writing, optional editing, formatting, optional SEO work, a quality gate, and publishing. The exported graph lives at langgraph_app.graph.workflow:main_graph, with langgraph.json declaring it for tooling.
This resembles an editorial handoff. A plan establishes the assignment, research supplies material, and the writer develops the argument. Editing and formatting operate on the result. Each handoff creates an opportunity to check whether the next stage has what it needs.
The repository preserves those boundaries. langgraph_app/graph/builder.py wires the graph. nodes.py contains contracted node runners, and workflow.py handles checkpoint integration. Agent implementations live under langgraph_app/agents/.
Consider a hypothetical debugging case: an article contains a relevant argument but uses unsuitable sources. A staged workflow lets you inspect the research output and the writer's input separately. You can then investigate whether retrieval returned poor material or whether the writer failed to use appropriate material already available.
Separate stages don't automatically guarantee better prose. They make defects easier to locate and give developers narrower places to apply validation. They also introduce more state and more failure points, so the architecture needs explicit contracts between stages.
Pydantic provides structured validation in the backend, while Zod serves validation needs in the frontend. These tools help enforce expected shapes. They cannot establish that a source supports a claim or that an article's argument is sound. Those require separate checks.
The backend's core/ directory includes contracts, verification, provenance, evidence, and assurance components. Their architectural role is to keep concerns such as source handling and validation available beyond an individual writing prompt. A prompt can request compliance; application code can enforce a boundary and record what happened.
Once generation becomes a workflow, the next dependency is admission control. The service must decide whether the request should enter that workflow at all.
The browser authenticates through NextAuth, with Google OAuth or credentials and support for TOTP multi-factor authentication and backup codes. The Next.js server then proxies requests to FastAPI. It sends an X-User-ID header alongside an internal secret.
An identity header is only trustworthy when the backend can establish who supplied it. A browser could otherwise claim another user's identifier. WriterzRoom therefore requires the shared internal secret at the backend boundary and refuses to start in production without it.
User-facing backend routes use require_user_id; internal endpoints validate the internal secret; administrative routes use require_admin_user. Those are separate checks because an internal scheduler invocation, an authenticated user request, and an administrative action carry different permissions.
Admission also checks entitlement, rate limits, and the submitted work request and sources. Credits are deducted in a database transaction before generation begins. That ordering prevents expensive work from starting before the service has established that the account can fund it.
It creates an obligation, too. If the workflow fails after deduction, the system needs a reliable way to resolve the charge. The later recovery layer follows directly from this decision.
The generation API separates submission from completion. POST /api/generate validates the request, deducts credits, and returns a request_id with status and stream links. The graph continues as an asynchronous background task inside the receiving instance.
Progress has two different paths. Each stage updates the generation's database row, and GET /api/generate/status/{id} reads that durable record. Any backend instance can answer the status request.
Live events travel through the in-process broker at core/stream_broker.py. Those events are instance-local. If a later connection reaches another instance, that instance does not share the original broker's memory.
A hypothetical frontend should therefore treat streaming as a responsive display channel while retaining polling as a recovery path. A disconnected stream should trigger reconciliation against stored status. It should not, by itself, turn a running generation into a failed one.
Graph checkpoints add another persistence mechanism. AsyncPostgresSaver stores workflow checkpoints and supports interruptions such as Premium human review. Stored progress answers where the workflow stands; a checkpoint preserves the information needed for workflow continuation. Neither makes an interrupted process resume automatically without coordinating code.
With workflow state and admission established, external integrations become clearer boundaries.
WriterzRoom's model clients connect to Anthropic, OpenAI, and Gemini through Vertex AI. Research includes Tavily web search, academic retrieval, curated and private corpora, and live connectors. Voyage AI provides embeddings and reranking, while PostgreSQL with pgvector supports vector search.
These components do different jobs. Search finds candidates. Embeddings support comparisons based on meaning. Reranking reorders retrieved material for relevance. The writing stage still needs rules about which material it may use and how that material supports the requested output.
A hypothetical request to write from uploaded internal documents illustrates the consequence. The product needs to preserve the distinction between private material and general web results throughout retrieval and drafting. Combining everything into one undifferentiated prompt would make those boundaries harder to inspect.
File handling adds another layer. The backend parses PDFs, Word documents, spreadsheets, and strict UTF-8 text formats. Parsing extracts content; it does not establish the content's accuracy, permissions, or suitability for publication.
Publishing follows the same boundary-based design. Adapters exist for destinations including Medium, WordPress, Ghost, dev.to, Hashnode, LinkedIn, Beehiiv, Notion, and Webflow. Each destination has its own credentials and failure behavior. A successful generation and a successful publication are separate outcomes.
That separation lets the service retain generated content even when a destination is unavailable. It also creates room for destination-specific formatting without forcing the writer to encode every platform's requirements into the draft.
Specialized tools run separately, too. A Model Context Protocol service exposes legal verification functions, including citation parsing and checks involving reporters, courts, and years. A separate curated-corpus service exposes read-only search and listing capabilities restricted to curated rows.
These service boundaries narrow access as well as organize code. The curated-corpus tool's read-only scope prevents a search operation from becoming a write operation. The internal automation repositories are also separate from customer generation nodes: analytics, prospect research, outreach drafting, and internal content planning have their own deployments and responsibilities.
The next documented changes address a mismatch between application logic and process lifetime.
The current deployment uses Cloud Run for the frontend, backend, specialized tool services, and a database migration job. Cloud SQL supplies PostgreSQL with pgvector. The database uses private networking, regional high availability, backups, and point-in-time recovery.
Terraform defines infrastructure in terraform-main.tf, but infrastructure changes are applied deliberately rather than automatically by CI. Application delivery follows a separate pipeline through GitHub Actions and Cloud Build.
Secret Manager stores credentials with per-secret access bindings. Gemini uses the backend service account's Vertex AI permissions through Application Default Credentials, which avoids a separate API key for that integration. Other external boundaries retain their own credential requirements.
Keeping a minimum instance available reduces cold-start exposure, but it also creates a baseline cost. More importantly, an available instance is still replaceable. A deployment or failure can remove the process holding a generation.
This limitation produced concrete changes. Earlier email automation relied on in-process asyncio.sleep. The process ended before the waiting task could send its message. Scheduled jobs now trigger onboarding and reminder work through internal endpoints.
Generation recovery addresses the same lifetime problem with a financial consequence. Credits are deducted before work starts, and the workflow's exception handler contains refund logic. If the instance dies, that handler cannot run. The database can retain a "running" generation whose executing process no longer exists.
The generation sweep identifies stale work, marks it failed, and refunds it. A future worker queue would still need a recovery mechanism, because a worker can claim a job and die before recording completion.
Other scheduled jobs handle publishing, source validity checks, regulatory monitoring, review assignments, and promotional expiration. Their schedules avoid placing every network-bound batch in the same time slot. The validity scan also runs less frequently than short-interval operational work because it contacts third-party hosts.
Every invocation is recorded in ScheduledJobRun, and /admin/jobs exposes operational states such as healthy, failing, overdue, or having no recorded runs. A configured schedule alone cannot tell you whether its work completed. The run record supplies that missing evidence.
Reliability also depends on how changes enter production.
The CI pipeline checks frontend linting, types, tests, and the production build. It runs backend tests, documentation validation, specialized tool tests, migration replay, and security scans. The tool tests include verification evaluation gates designed to test both detection and false positives.
Deployment waits for the blocking jobs and runs on pushes to main. Production deployments are serialized so separate merges cannot race through a partly applied release. There is no staging environment, so merging to the main branch changes production.
That arrangement raises the value of disposable migration tests and compatible changes. Cloud Build applies database migrations before updating either application service. If migration fails, the release stops before those updates.
Cloud Run revisions allow application rollback by shifting traffic to a previous revision. Database migrations are forward-only. A previous application revision must therefore remain compatible with the schema it encounters, or the team needs a compensating migration.
A hypothetical field replacement makes this concrete. Adding a replacement field, supporting both representations, migrating existing data, and removing the old field later gives older code a compatibility window. Removing the original field immediately can leave rollback code unable to operate.
Capacity exposes another non-obvious constraint. Increasing backend instances also increases database connection pools. Compute may have room to expand while PostgreSQL runs out of available connections. The effective ceiling comes from the combined system.
The documented adjustment is an explicit Prisma connection limit, but it cannot simply be added to the shared database URL. That URL also serves the checkpoint driver and health checks, which do not accept the Prisma-specific parameter. Separate client-compatible connection strings are needed first.
Memory is another boundary because concurrent generations share instance resources. Request concurrency should account for long-lived workflow state, rather than borrowing settings suitable for short HTTP handlers.
scripts/loadtest.py separates read testing from generation testing. The generation profile invokes paid work and requires explicit cost acknowledgment. A mocked generation mode was removed to avoid carrying test-only branches in the customer execution path.
For someone using WriterzRoom, architecture becomes visible through behavior: progress survives a page refresh, review can interrupt a workflow, generated content remains available after a publishing failure, and abandoned paid work has a recovery path.
These behaviors depend on the layers working together. A review screen needs persistent workflow state. A credit display needs a ledger that agrees with generation outcomes. A publishing status needs to distinguish content creation from delivery to an external platform.
The frontend stack supports those interactions through a rich-text editor, reusable interface components, server-state management, and client-side validation. TipTap provides editing capabilities, while TanStack React Query helps manage fetched application state. The backend remains responsible for authoritative decisions about permissions and work status.
For developers integrating or testing the service, the submission-and-status contract is more useful than assuming a request returns a finished article. Preserve the returned request identifier, use the provided links, and handle connection loss independently from generation failure.
A server-side status request follows this shape:
curl "$BACKEND_URL/api/generate/status/$REQUEST_ID" \
-H "X-User-ID: $USER_ID" \
-H "X-Internal-Secret: $INTERNAL_API_SECRET"
That internal secret belongs in a trusted server environment. It must never be embedded in browser code. If you test the backend directly in Postman, keep the secret in a private environment and keep exported collections free of credentials.
The repository includes no Python SDK. Integrators should use the actual HTTP contract rather than assume a supported client library exists.
The same practical discipline applies inside the product. Inspect stored stage progress before investigating prose. Check publication status separately from generation status. Correlate errors through structured logs, Sentry, and LangSmith, while avoiding unnecessary exposure of manuscript content.
WriterzRoom's current architecture creates a path toward moving execution beyond the lifetime of the request-receiving instance. Durable job dispatch and a shared event channel could make work and live progress less dependent on process placement. They are prospective changes. They do not describe the current implementation.
That path is easier because the system already separates admission, graph execution, persistence, and recovery. Moving execution would not require inventing those responsibilities from scratch. It would require carrying their guarantees across another boundary.
The same applies to expanding review and publication capabilities. Explicit stages provide places to pause, inspect, and apply destination-specific rules. Persistent records provide something to reconcile when execution stops.
The useful next step is to strengthen those handoffs. A writing service becomes dependable when it can explain what it accepted, what it completed, what remains unresolved, and how it will recover. Better prose still matters. Architecture gives that prose a reliable route from request to publication.
WriterzRoom Team · October 4, 2026
Discussion
Share a question or perspective on this article. Comments are reviewed before appearing here.
Sign in to leave a comment.
Loading comments…