Every CRM on the market is a database that asks a salesperson to be its data-entry clerk. ORBIT inverts that. The system records, thinks, and drafts; every item arrives framed and ready; the human approves it or edits it and marks it ready.
This document specifies the product in full, on its own codebase, sharing nothing with anything else we run. Chapter 16 prices it honestly in people and time, and shows what the date does at different staffing levels.
What we are building, why it can win, and what it costs to build properly.
Whoever makes the CRM fill itself in — and then tells the rep exactly what to do next, already prepared — takes the category.
Not because the AI is smarter. Because the human effort required drops to near zero, and adoption is the only metric that has ever mattered in CRM.
CRM data entry has never been solved. Every generation shares one failure mode: the record is only as good as the rep's willingness to update it, and reps do not want to update it. Forecasting, reporting, coaching and handoffs all sit on a foundation the business knows is unreliable.
Capture is now cheap. Transcription costs fractions of a cent per minute. Email and calendar APIs are mature. Meeting-bot infrastructure is buyable. Recording every interaction a sales team has is an engineering problem now, not a cost problem.
Reasoning over that capture is now possible. With a properly built memory layer, a model can answer why is this deal at risk from evidence rather than from a rep's half-remembered notes.
ORBIT is an independent product on its own codebase. It shares nothing with anything else we run — not code, not infrastructure, not a database. It receives leads and engagement data from whatever the customer already uses at the top of the funnel, and turns everything after that first touch into a system that runs itself.
Every upstream tool is a provider — a peer entry in a registry, integrated the same way, privileged in no way. OXO is one of them, alongside Smartlead, Apollo, Outreach, a form, a CSV or a raw webhook.
The hardest thing about selling a CRM is that the customer already has one, and replacing it is a six-month project nobody wants to own. So ORBIT ships in two modes: as the system of record, or as a capture-and-intelligence layer that writes back into whatever they already run.
Mode B is the commercial wedge — it asks the customer to change nothing, produces the capture data that makes our AI good, and turns the eventual migration into a switch flip because by then we hold the entire history. Mode A is what that migration lands on, and it has to exist and be good on the day the first customer asks, or the wedge is a dead end rather than a door.
ORBIT is a separate codebase with separate infrastructure. Nothing carries over as software. What carries over is narrower and worth stating precisely, because overstating it would set the wrong expectation with this room:
| What transfers | What does not |
|---|---|
| Knowledge of how mail capture at scale actually fails — quota behaviour, push reliability, backfill shape | Any of that code |
| Vendor evaluation results for transcription, telephony and enrichment | The integrations themselves |
| Hard-won patterns for multi-tenancy, job orchestration and model cost control | The implementations |
| An understanding of the buyer, learned from selling to them | Distribution — every customer is a new sale |
| Relationships that make design-partner recruitment faster | Any enrichment dataset. ORBIT buys enrichment, and the margin model reflects that |
Engineers who have already made these mistakes once will move considerably faster than engineers who have not — but that is experience, not inventory, and the estimates in chapter 16 do not assume a head start.
This document specifies the product in full: six modules at depth, both modes, native mobile and a complete web application, the provider framework, the memory layer, the AI backbone and a separated analytics plane.
Built properly, that is a team of roughly twenty-two people and about eleven months to a commercially complete v1, with the first sellable release at month four.
Chapter 16 sets out the team shape, the phasing and a sensitivity table showing what the date does at six, twelve, twenty and twenty-two engineers. The scope in this document does not change across those rows — only the calendar does. Which row we operate on is a staffing decision, and it belongs to this room rather than to the specification.
Approve the scope, approve the team shape, and pick a row on the sensitivity table. Chapter 17.
Why every CRM rots, and the specific user we are going to design for without apology.
A CRM is sold to a VP of Sales as a system of record. It is experienced by a rep as unpaid administrative labour with no personal payoff. The rep does the typing; the manager gets the dashboard. That mismatch is the root cause of every CRM problem, and no vendor has fixed it — they have only added more fields.
We should be blunt about this internally, because it is the most important constraint in the document.
The target user will not read documentation, will not configure anything, and will not remember where a feature lives. They open the app, and if it does not immediately tell them what to do, they close it and go back to their inbox.
What follows: if a task requires the user to know what they want to do, we have already failed. This is not a niche concern — it describes the median salesperson, and it describes every salesperson on a Tuesday afternoon between two calls. Building for this user is the only genuine differentiator available in a category where everyone has feature parity.
| Conventional CRM asks the user to | ORBIT instead |
|---|---|
| Remember to log a call | Records, transcribes and summarises it |
| Decide who to follow up with today | Presents seven ranked items, already prepared |
| Write the follow-up email | Drafts it from what was actually said, marked ready to approve |
| Update the deal stage | Detects the change and asks for one confirmation |
| Remember what was promised | Extracts commitments into dated tasks |
| Search for context before a call | Pushes a prep brief thirty minutes ahead |
| Build a report | Answers the question, with citations |
| Learn a workflow builder | Takes the automation in plain English |
Four groups of competitor, an honest read of the closest one, and the four wedges we can hold.
| Group | Examples | Structural weakness |
|---|---|---|
| Legacy system of record | Salesforce, HubSpot, Dynamics | AI bolted onto a schema designed in 2004; still needs manual input; desktop-shaped |
| Modern but manual | Attio, Folk, Twenty | Beautiful and fast, but still a database the human must feed |
| Intelligence layers | Gong, Clari, Avoma | See conversations but don't own the record; an expensive add-on to an expensive CRM |
| AI-native revenue OS | Reevo, Day.ai, Attention | Desktop-first; tied to their own funnel; still assumes an engaged user |
Reevo is the sharpest expression of the thesis we are pursuing and should be treated as the benchmark. They are an AI-native revenue operating system combining native CRM, prospecting, engagement and deal intelligence, keeping the whole sales workflow inside one system.
Their core insight, which we should steal outright. No single tool sees the whole deal — a CRM nobody updates, a call recorder, an outbound tool — so each sees a fragment, and AI bolted onto a fragment is guessing. Their answer is that every call, email and deal feeds one shared memory, so the system drafts follow-ups, updates records and flags risk on its own. They frame that accumulated context as the most valuable thing a sales team creates.
Their resources. $80M from Khosla Ventures and Kleiner Perkins, explicitly to rebuild the GTM system from the ground up rather than bolt lightweight AI onto an existing product. SOC 2 compliant, ISO 27001 certified, three custom-quoted tiers.
Honest assessment. Well funded, credible, and roughly 18 to 24 months ahead on product surface. We will not beat them by matching feature for feature on their timeline. We beat them on four axes they cannot easily take, two of which conflict with their own positioning.
They sell a platform that replaces your funnel tools. We plug into whichever ones you already pay for. They cannot copy this without undermining their own pitch — every provider we support is a product they are trying to displace. A positioning conflict, not an engineering gap, which makes it durable.
Their pitch is retire your sales stack — a large, slow, risky purchase. Ours is change nothing, and we will make your existing CRM correct. An easier first sale and a shorter cycle, converting to full replacement later from a position of holding all the data.
Everyone else assumes an engaged professional who opens an app and works it. We assume the opposite: a hard cap on daily items, everything pre-prepared, no configuration to get value. Hard to retrofit onto a broad enterprise surface, because enterprise asks for options and options are what we are removing.
The category is desktop-shaped because it was built for the manager. We design the rep's loop for a phone first and ship it natively. A rep who clears their day in ninety seconds on a phone uses the CRM daily; one who must open a laptop does not.
Ten rules, each written so it can be used to reject a design in review.
No user should ever type something the system could have observed.
Any feature requiring the rep to enter data that exists in an email, a call or a calendar is a bug. Measured by zero-input ratio — the share of record changes originating from capture rather than a keystroke. Target 85 percent.
There is no non-AI path through this product.
Capture, extraction, memory, ranking, drafting, routing and narration all run through one orchestration layer with one registry. A feature that works only when AI is switched off does not belong here.
The system prepares; the human approves, or edits and marks ready.
No blank compose box, no empty form, no unscheduled slot. Every item the user meets is already framed and one gesture from done. The edit path is not a failure — the diff between our draft and what the human sent is the most valuable signal we collect.
"What should I do right now?" is the only question the home screen answers.
Not a dashboard. Not a record list. A ranked, finite, executable queue. If a screen presents more than one job to be done, it is not the home screen.
No integration is special-cased, ever.
Mail, calendar, funnel tools, CRMs, meeting bots and raw webhooks all arrive through one framework with one contract and one normaliser. The moment a source gets bespoke handling in application code, the cost of the next source doubles, and provider count is a direct commercial input.
Nothing is overwritten, and nothing is consulted selectively.
Append-only: new knowledge supersedes old without destroying it, so the system can always answer what changed and when. And every AI task reads from it — there is no path that generates without first consulting what the workspace knows.
Every AI claim carries a citation to its source artifact.
"This deal is at risk" is unusable. "The economic buyer has not been on a call in 24 days, and on 3 March the champion said budget moved to Q3 — play from 14:02" is actionable and falsifiable. Enforced at the API layer, not by prompt instruction.
No user-facing screen may wait on a model call to render.
Everything seen on open is precomputed. AI runs on ingest and on schedule, not on page load. A model outage degrades freshness; it must never produce a blank screen.
The system starts by asking permission for everything.
Three autonomy tiers, configurable per workspace and per task class, progressing as measured accuracy justifies it. One badly auto-sent email costs more trust than fifty good suggestions earn.
A missing channel makes every downstream feature lie.
Capture completeness outranks any intelligence built on top of it. A memory layer holding seventy percent of the interactions is not seventy percent as good — it is actively misleading, because it will confidently report that nothing happened on an account where plenty happened somewhere we cannot see. When a capture gap and a feature compete for the same sprint, the gap wins.
Eight components. Almost all of the engineering value and all of the defensibility sits here, beneath the visible product.
The six modules are the visible surface, and they are thin. Build the platform correctly and modules become weeks of work each. Build modules first and retrofit the platform, and we ship a worse Pipedrive.
One immutable, append-only, normalised event log per workspace. Every provider writes here first; memory, actions, timeline, workflows and reporting all read from it.
workspace_id so ordering is guaranteed per tenant and one heavy workspace cannot starve the rest. Replay is the property that matters most — when extraction improves, we reprocess history rather than only benefiting on new data.occurred_at and ingested_at separately, and every derived artifact is recomputable rather than incrementally-only.The full lead lifecycle — first touch, every email and call, every meeting, conversion, client data, churn — is a query against the spine filtered by entity. Building it as a log rather than as a UI feature is what makes lifecycle, reporting, memory and audit all fall out of one structure.
Row-level, with a workspace_id on every table, enforced by Postgres row-level security. The alternative — a schema per tenant — is defensible right up until you notice that neither the vector index nor the analytics store has a schema equivalent, at which point the product has two different tenancy models and the seam between them is where cross-tenant leaks are born.
| Consideration | Schema per tenant | Row-level with RLS |
|---|---|---|
| Migrations at 1,000+ tenants | N schemas by M migrations, long lock windows | One migration |
| Cross-tenant operational analytics | Union across schemas | Native |
| Vector and ClickHouse alignment | No schema equivalent — the model diverges | One partition key everywhere |
| Blast radius of a query bug | Structurally isolated | Depends on RLS being correct, so it must be tested |
| Per-tenant residency | Easier | Requires regional sharding |
Guarded by a mandatory CI test that fails any query lacking a workspace predicate, with physical database separation offered as an enterprise tier for customers who require it contractually.
One Person, no separate Lead object. The Lead-versus-Contact split is a historical artifact that forces a lossy convert operation and confuses users. We model a single person with a lifecycle_stage, and map at the sync boundary when a downstream CRM needs the distinction.
Custom fields, typed and described. Every CRM dies without them, but unlimited custom fields destroy the AI layer — a model cannot reason over four hundred arbitrary columns. Typed fields with mandatory descriptions, where the description is fed to the model so it knows what the field means. A field the AI cannot understand is a field it cannot fill in, and tenet one says it must.
The unglamorous component that decides whether the product feels magical or broken. If an email address, a name on a Zoom call and a phone number in a dialer log are three different people in our system, the timeline fragments and every AI output is wrong.
Deterministic match first — email, E.164 phone, calendar attendee ID, provider external ID — then domain attribution with a maintained exclusion list of free and ISP domains, then probabilistic match on name and company, then model adjudication for the residual ambiguous set only.
Roles across owner, admin, manager, rep and read-only, with record visibility open, team-scoped or owner-scoped per workspace. The part that is easy to get wrong:
The permission filter runs on the retrieval index, not on the output. Filtering after generation means the model already saw data the user is not allowed to see, and will leak it in paraphrase.
Alongside that: jurisdiction-aware recording consent shipped with the recording feature rather than after it; per-workspace keys for recordings and transcripts, with bring-your-own-key at the enterprise tier; and an erasure path that propagates through raw payloads, activities, facts, entity cards, embeddings, the analytics plane and any synced downstream CRM. Designed on day one, because retrofitting deletion through a vector index and a columnar store is among the most expensive mistakes available to us.
For this user, notification strategy is the product — a user who mutes us is a churned user who has not cancelled yet. One scheduled digest a day plus at most three event-triggered pushes, enforced centrally so no module author can add just one more. Only three things earn an interrupt: a meeting starting soon with its brief attached, a prospect reply on an active deal, and a post-call summary that is one tap from sent. If a user consistently ignores a class, we stop sending it and fold it into the digest — measured and automatic, with no settings screen.
Native bi-directional Salesforce and HubSpot, with the long tail via a unified-API vendor. Field mappings auto-generated on connect by inspecting the customer's schema, because our user will not configure a mapping screen. Field-level conflict policy with a conservative default: ORBIT writes activities, summaries and enrichment freely, and proposes but never silently overwrites a human-edited deal field. Overwriting a rep's close date on day one is how we get uninstalled.
Backfill on connect is what produces the first ten minutes of value: the customer connects their CRM and immediately sees a populated, intelligent timeline of a deal they already know — proof the system works, using their own data.
Webhook or poll is the wrong question. This chapter argues for both, always, and specifies the machinery that makes that affordable for two engineers.
Every source of data — a funnel tool, a mailbox, a calendar, a meeting bot, a CRM, a raw webhook — is a provider. Providers are declared, not coded. This is the single highest-leverage piece of architecture in the product, because the count of providers we support is a direct commercial input and a direct engineering cost, and the framework is what decouples the two.
Most systems treat webhooks and polling as alternatives: use webhooks if the provider offers them, fall back to polling if not. That is wrong, and it fails in production in a way that is very hard to debug later.
Webhooks give you latency. Polling gives you truth. A provider integration that has only one of them is either slow or quietly lossy — and quietly lossy is worse, because the memory layer will confidently report that nothing happened.
Webhooks are best-effort by construction. Deliveries are missed during our own deploys and outages. Providers drop them under load, cap retries, silently disable endpoints after consecutive failures, and occasionally ship bugs that stop emitting an event type entirely. None of those failures announce themselves. Polling with a cursor, by contrast, is self-healing: whatever was missed is picked up on the next sweep, because the cursor never advanced past it.
So the recommendation is a hybrid by default: webhooks for freshness where a provider supports them, plus a lower-frequency reconciliation poll that exists purely to catch what the webhooks lost. Because every write is idempotent, the overlap costs nothing but a few wasted comparisons.
With two engineers, the cost of the tenth provider matters more than the cost of the first. So a provider is declared as data — capabilities, auth, endpoints, rate limits, pagination style — plus one small mapper function per resource type. The framework does everything else.
| Manifest field | What it declares |
|---|---|
| auth | OAuth2, API key, or basic — with the token-refresh strategy and the scopes required |
| supports_webhook | Whether a push path exists, its signature scheme, and which event types it emits |
| supports_poll | Whether a pull path exists, and the cursor style: opaque token, delta link, or updated_since |
| cursor_field | Which field is monotonic, and whether it is server-assigned or client-clock dependent |
| pagination | Cursor, offset, or link-header, and the maximum page size |
| rate_limit | Requests per window, whether the quota is per app or per connection, and how 429s report retry timing |
| event_id_field | The provider's own event identifier, if any — otherwise the framework hashes the canonical payload |
| supports_backfill | Whether history is reachable, and how far |
| resources | The resource types to sync, each with a mapper to the envelope |
Target: a new provider is a manifest, a mapper and a fixture-based test — two days of work, not two weeks. Everything else in this chapter exists to make that number true.
/hooks/:provider/:connection_token. The token identifies the connection without a database lookup on an unauthenticated path.occurred_at in the payload, and from a version or etag per entity where the provider supplies one. A stale update must lose to a newer one already applied.| Provider offers | What we run | Reasoning |
|---|---|---|
| Webhooks and a cursor API | Both — push for latency, hourly sweep for truth | The default. The sweep is cheap and catches everything push loses. |
| Webhooks only, no queryable history | Push, plus durable raw retention and gap alarms | We cannot self-heal, so we must at least detect loss and tell the customer |
| Cursor API only | Adaptive polling | Latency is bounded by the interval; tighten it for connections that matter |
| Neither — full-list only | Snapshot diffing on a slow schedule | Expensive and last-resort. Cap the object count and be explicit with the customer about freshness. |
| Push with guaranteed ordering and delivery | Push alone, sweep weekly | Rare. Still sweep, because the guarantee covers their side of the wire and not ours. |
Both paths are at-least-once. Exactly-once delivery does not exist across a network boundary, and pretending otherwise produces bugs that surface months later as duplicated activities. Instead we make every write idempotent and get effectively-once behaviour:
dedupe_key = provider : connection : external_id, or a canonical payload hash when the provider gives no stable identifier.Because the memory layer's credibility depends on completeness, absence of data has to be treated as a signal rather than as an absence of signal. Three mechanisms:
Gmail and Microsoft Graph are each a substantial integration to do properly — push channel renewal, quota behaviour, threading semantics, delta tokens, historical backfill, and the long tail of tenant-specific admin configurations. Doing both well is a meaningful fraction of a year for a dedicated engineer, and it is undifferentiated work: no customer chooses us because our Graph integration is elegant.
The recommendation is to buy a unified mail and calendar vendor to reach coverage early, and build the Google path natively in parallel once the product is proven. Because both are manifest entries, the switch is a configuration change and nothing above the normaliser moves. That substitutability is precisely why the framework is built first — it converts a strategic build-versus-buy argument into a reversible operational one.
One registry, one router, one ready-state model. Every intelligent behaviour in the product is declared here before it is written.
AI is not a layer in this system. It is the column every other layer calls through — and that only works if every task is declared, versioned and measurable rather than scattered as prompt strings across a codebase.
Every AI operation in ORBIT — "summarise a call", "extract commitments", "classify a reply", "draft a follow-up", "score deal risk", "narrate a report" — is a registered task. A task is a first-class database record, not a constant in application code.
| Property | Why it matters in practice |
|---|---|
| Provider and model per task | Model price-performance moves monthly. Cheap tiers handle extraction and classification; frontier models handle risk reasoning and the Ask interface. That split is the single largest lever on gross margin. |
| Versioned prompts | A prompt is an artifact with an owner, a changelog and an eval score — not a string edited in a pull request nobody reviews. |
| Staged rollout per task | Internal, then design partners, then ten percent, then all — with automatic rollback on a metric drop. Per task, not per release. |
| Output provenance | Every stored AI output records the task key, model and prompt version that produced it. Without this, "quality got worse last month" is unanswerable. |
| Per-workspace override | An enterprise customer requiring a specific provider, or a region requiring in-country inference, is a configuration change. |
| Cost attribution | Spend is measured per task and per workspace as a first-class metric, not reconstructed from provider invoices at quarter end. |
| Failover | Declared fallbacks per task mean a provider outage degrades quality slightly instead of taking a feature down. |
No AI call may be made outside the registry. A prompt string in application code fails code review. This is worth enforcing with a lint rule, because the discipline erodes quietly and the cost of restoring it later is a full audit of every AI behaviour in the product.
The second half of the backbone is a state machine that every AI-produced artifact moves through — a drafted email, a proposed stage change, a suggested meeting slot, an extracted commitment, a generated report narrative. This is what tenet three means concretely.
Why this is architectural rather than cosmetic. Making ready an explicit persisted state, rather than an implicit consequence of a user tapping a button, gives us four things at once: a single audit trail of who cleared what; a uniform undo boundary; a place to attach the autonomy tier per task class; and a clean measurement point for the metric that matters most — the share of drafts approved without edit.
Written forever, read always. The component that decides whether ORBIT is a differentiated product or a thin wrapper over a language model.
The obvious approach — embed every email and transcript, retrieve top-k on a query — produces a good demo and then fails in production, predictably:
The answer is layered memory, where each layer answers a different class of question, and a router that decides which layers to consult.
The memory layer is append-only. Nothing is ever overwritten or deleted in the course of normal operation — only superseded.
When a new fact contradicts an existing one, the old fact is marked superseded_by and stays in place. Retrieval defaults to current truth, but the history remains queryable. That single design choice gives us three things a mutable store cannot:
The only exception is a lawful erasure request, which propagates through every layer including embeddings and the analytics store. That path is designed on day one, because retrofitting deletion through a vector index and a columnar database is one of the most expensive mistakes available to us.
The complement of forever-written is always-read. No AI task generates without first consulting memory. There is no code path where a draft is written from the immediate context alone — the entity card is in the prompt, the relevant facts are retrieved, and the citation chain is attached, every time. A task that skips memory is a task that will eventually contradict something the customer told us, which is the fastest way to lose their trust.
| Field | Value | Why it exists |
|---|---|---|
| statement | Budget approval moved from Q2 to Q3 due to a hiring freeze | One claim, not a paragraph |
| source | call act_9912 at 14:02 | Every citation resolves to a playable moment |
| asserted_by | per_77, champion | A fact from the economic buyer outranks the same fact from a junior contact |
| confidence | 0.86 | Low-confidence facts are retrievable but never the sole basis for an automated action |
| valid_from | 2026-03-03 | Decay is tuned per predicate — budget decays fast, org structure slowly |
| superseded_by | null | The append-only mechanism. History survives; retrieval defaults to current truth |
| produced_by | extract.commitments v7 | The registry stamp, so a regression can be traced to a prompt version |
Recommendation: pgvector, in the same Postgres as the system of record. Phase-one volume is low millions of vectors, well within range, while our filtering needs — entity, date, permission — are complex. That is precisely pgvector's sweet spot, and it means one tenancy model instead of two, with facts and their embeddings written atomically. Migrate behind an interface later if volume demands. Do not add a stateful service in month one to solve a year-three problem.
L4 is strictly within a workspace. Cross-customer learning is commercially tempting and contractually radioactive. If we ever pursue it, it is opt-in, aggregated, differentially private, and a separate board-level decision.
The user does not know what to do, so the platform decides — and has the work already done before they look.
This is the product. If a user opens ORBIT and sees a small, correct, ranked list of things to do, each already prepared, we win. Everything else is supporting infrastructure.
A system-generated, ranked, expiring item with a precomputed payload, sitting in the drafted state and one gesture from ready. Four properties, checked at generation:
States its reason with a citation. Never "follow up with Nexus" — always "Nexus has gone quiet for twelve days and closes in nine."
The work is already done. Draft written, slot chosen, number dialled. Never a suggestion attached to an empty compose box.
Approve, or edit and mark ready, in a single pass of a few seconds.
Carries a TTL. A stale item is worse than none — it teaches the user the queue is noise.
A nightly batch regenerates the queue per workspace in local time before working hours, using the batch API at reduced cost. Event triggers — a reply arrives, a call ends — recompute affected items within about sixty seconds. Re-ranking as the day passes is cheap and deterministic, with no model call at all.
| Signal | What it means | What we do with it |
|---|---|---|
| Approved unedited | Item and payload both right | Strongest positive reinforcement |
| Edited, then marked ready | Right item, wrong draft | The highest-value signal we collect. The diff feeds the eval set of the registry task that produced the draft |
| Snoozed | Right item, wrong time | Adjust the timing model |
| Dismissed | Wrong item | Suppress that pattern for this user and account |
If any action type drops below thirty percent acceptance, it is automatically demoted out of the default queue pending review. The queue polices itself; we do not wait for a customer to complain.
A separate store, a separate deployment, and an hourly spool. Reporting and analytics never touch the transactional database.
This is an architectural separation, not a module boundary. The application that serves the rep's queue and the system that answers "what is our win rate by segment this quarter" have opposite access patterns, opposite failure tolerances and opposite scaling curves. Coupling them means a manager running a heavy report degrades a rep's screen — and that is exactly the failure that gets a product uninstalled.
Both Postgres and ClickHouse could carry this. The recommendation is ClickHouse, for four reasons specific to us:
The trade-off, stated honestly: a second store means a second tenancy enforcement point, a second migration story, and eventual consistency between what the record says and what the report says. The hourly cadence and the visible as-of stamp are how we make that consistency gap explicit rather than confusing.
Real-time analytics is a tempting default and the wrong choice here. Three arguments:
A figure that shifts mid-conversation is a figure nobody trusts. An hourly boundary means a pipeline review runs on one consistent snapshot for its duration.
Incremental view refresh once an hour is a fraction of the cost of maintaining continuously consistent aggregates, and nobody makes a different decision because a number is fifty minutes old.
Activities arrive out of order — a backfill, a delayed sync. An hourly window lets late events land before aggregates are computed, rather than producing figures that silently revise.
The rep's own queue and record are served from the transactional plane and are always current. Only aggregate reporting runs on the hourly cadence — and an on-demand refresh is available for the moment before a board meeting.
Between the data and every consumer — dashboards, scheduled reports, and the plain-language Ask interface — sits a governed metric layer. Every metric is defined once, in code, versioned and tested: pipeline coverage, win rate, cycle time, stage conversion, quota attainment.
The naive alternative is text-to-SQL over the raw schema. That approach hallucinates joins, silently miscounts on one-to-many relationships, and produces different answers to the same question asked twice. Text-to-metric over a governed layer is a bounded, verifiable problem — the model selects metrics, dimensions and filters from a finite validated set, and never authors SQL. This one decision is the difference between a reporting module a CFO trusts and one they do not.
If AI retroactively corrects a stage-entry date, last month's conversion report changes. This needs as-of semantics: stage history is immutable and append-only, and every report states the point in time it reflects. Easy to design in now; extremely expensive to retrofit.
| Consumer | Reads from | Freshness |
|---|---|---|
| Scheduled reports and dashboards | Materialised views | Hourly, stamped |
| Ask Anything, aggregate questions | Semantic layer over MVs | Hourly, stamped |
| Ask Anything, entity questions | Memory layer, transactional plane | Live |
| Proactive insight pass | Materialised views | Daily and weekly |
| Predictive model training and scoring | Feature tables built in the same spool | Daily |
| Today queue and record screens | Postgres | Live — never routed here |
Lead management, pipeline, meetings, workflows, reporting, analytics — each with its AI behaviours, its hard problems, and what it deliberately does not do.
Take a person from exists to qualified and in pipeline with the minimum possible rep decision-making, and never let a lead go cold through inattention.
Leads arrive from any connected provider and are treated identically once past the ingestion contract. Auto-enrichment on arrival from a bought enrichment provider plus web signals means the rep never sees an empty record — title, seniority, department, company size, funding, tech stack, all present before the rep looks.
Qualification with reasons. A score alone is useless. "Matches your ICP on size and industry; the title is one level below your typical buyer; they opened the pricing page twice this week" is actionable. The ICP model is learned from the workspace's own closed-won history, not configured in a form.
Intent-classified inbound. Replies classified as interested, objection, not now, wrong person, unsubscribe or out-of-office, each with the right next item pre-drafted and waiting in the queue. "Not now" auto-schedules a dated revisit rather than dying.
Decay detection. A lead past a learned inactivity threshold generates a re-engagement item with a drafted message — or an explicit one-tap disqualify, because a clean pipeline is worth more than a large one.
ORBIT does not build a campaign engine. Sequencing at scale belongs to whatever the customer already runs, and duplicating it would put us in competition with our own integration partners. We do light one-to-one and small-batch follow-up; everything larger flows in from the connected source.
Make the pipeline reflect reality without a rep maintaining it, and make risk visible before the deal is lost rather than at the quarterly review.
Self-updating records. Stage changes, close-date shifts and next steps are drafted from conversation content, each with a citation, and sit in the queue awaiting approval. High-confidence, low-risk updates may move to ready automatically with a visible undo once a workspace has earned that autonomy tier.
| Risk signal | What it detects |
|---|---|
| Silence | Days since meaningful contact, measured against this deal's own baseline rhythm |
| Single-threading | A deal with one engaged contact is a deal with one point of failure |
| No economic buyer | Historically the strongest predictor of a deal that dies late |
| Sentiment trajectory | Direction across recent interactions, not absolute value |
| No scheduled next step | The most predictive indicator, and the most fixable |
| Slipped close dates | Count and magnitude, versus this workspace's historical stage velocity |
Relationship mapping builds an inferred org chart per account from thread participation, meeting attendance and reply patterns — surfacing the gap that matters: you have no relationship with the person who signs.
Evidence-based forecasting. Commit, best case and pipeline derived from signals, presented alongside the rep's own call with the delta explained. We do not replace the rep's judgment; we make the disagreement visible and reasoned, which is what a sales leader actually wants.
A wrong auto-advance corrupts the forecast and destroys credibility with the sales leader — who is the economic buyer of this product. Every workspace starts at propose-only and graduates per task class on measured accuracy. Our own forecast accuracy is tracked and published inside the product, because our number will sometimes disagree with the VP's and it has to survive that meeting.
Walk into every meeting fully prepared without preparing; walk out with the record updated, commitments captured and the follow-up ready to send.
Extraction priorities, in order: commitments with dates and owners — the highest-value extraction in the entire system — then objections and competitor mentions, sentiment trajectory, talk-ratio for coaching, and deal-field proposals. The follow-up is drafted in the rep's own voice, learned from their sent mail.
iOS does not permit native call recording. A platform constraint, not an engineering gap. Options and the recommendation are in chapter 12.
Transcription on Indian, accented and code-switched speech. A real accuracy risk for our likely first market. Benchmark vendors on our audio, not their marketing benchmarks, and boost vocabulary with account and product names already in the CRM.
Bot presence is socially awkward. "Notetaker has joined" changes the meeting. Needs tasteful naming, a per-workspace policy, and a no-bot mode that falls back to post-call voice capture.
Let a team encode how we sell here once, and have it run reliably — without anyone learning a workflow builder.
Visual builders are powerful and used by approximately nobody outside RevOps. But letting the model decide what to do at runtime is unacceptable for a system that emails customers and mutates records: non-deterministic, unauditable, untestable. The resolution is the key architectural line in this module.
The model authors and explains the workflow. Deterministic code executes it. The model runs at execution time only inside explicitly marked content-generation steps, and each of those is a registered task.
Optimise for the template gallery, not the authoring canvas. Ship twenty to thirty curated templates — stalled-deal nudge, post-demo follow-up, no-show recovery, champion-departure alert, renewal at T-90, unworked-lead SLA. Most customers will never author a workflow; they will activate three templates.
Deliver numbers a sales leader will defend in a board meeting. Correctness and consistency matter more than flexibility — which is why reporting runs entirely on the analytics plane described in chapter 10, against governed metrics, on the hourly cadence, with a visible as-of stamp.
The library covers pipeline, forecast, activity, conversion funnel, rep leaderboard, source attribution and cohort retention, with scheduled delivery to email and Slack carrying an AI-written narrative alongside the numbers, and drill-through from any figure to the underlying records and ultimately to the source call or email.
On mobile, reporting is a manager surface and shows three headline numbers, the direction of travel, and what changed and why in one sentence with a link to the evidence. A briefing, not a shrunken dashboard. On the web it is the full instrument.
Answer questions nobody thought to ask, and surface things nobody thought to look for.
Plain-language questions across memory and the semantic layer, in the app, on the web, in Slack, or by replying to the digest. Every answer carries citations and, where relevant, a drafted item in the same response — the answer and the act are not separate journeys.
The lazy user does not ask questions, so a scheduled pass pushes what it finds: conversion anomalies, cohort divergence, emerging objection themes, stage-velocity decay, at-risk customer clusters. This matters more than capability one.
Calibrated win probability, predicted close date against the rep-entered one, churn risk, pipeline sufficiency. Gradient-boosted trees on features built during the hourly spool — cheaper, faster and explainable. The model predicts; the language model only narrates.
A five-rep team closing twenty deals a quarter cannot support a per-workspace model. Hierarchical models with sensible priors, visible confidence intervals, and a willingness to say not enough data yet — a correct and trust-building answer.
Campaign management and outbound sequencing, which belong to the connected source tools. CPQ, quoting and e-signature. Support ticketing. ERP and finance integration. Partner and channel management. Field service. Commission management. Stated plainly so nobody assumes otherwise.
Mobile-first, not mobile-only. The rep's daily loop is designed for a phone; the manager's work is designed for a browser. Neither is a shrunken version of the other.
Mobile-first is a statement about sequence and priority, not about scope. It means the rep's daily loop is designed for a phone before anything else is designed at all — because that is the loop that decides whether the product is adopted. It does not mean the web is an afterthought. The web is where the depth lives, and a large part of the buying decision happens there.
The rep's entire day in ORBIT should be completable in under three minutes, standing up, one-handed, without typing.
If that sentence is true, the product wins on adoption regardless of feature parity. Every mobile design review asks whether a change moves us toward it or away from it.
| Tab | The job it does |
|---|---|
| Today | "What do I do now?" — where most sessions begin and end |
| Deals | "Where does everything stand?" — a risk-ranked vertical list, not a kanban, which is a desktop metaphor |
| Search / Ask | One field. It searches records and answers questions. Not two features. |
| Me | Profile, settings, targets. Deliberately shallow. |
Six gestures, consistent everywhere, no exceptions: swipe right to approve, swipe left to snooze, long press to dismiss, tap to expand with the evidence, hold the mic to speak, pull down to refresh. Consistency is what lets a disengaged user become fluent without ever learning anything.
Typing on a phone is the friction that kills mobile CRM, so typing is not the input method. Hold the mic and say: "just got off with Meridian, they want a security review before signing, Priya is bringing in their CISO next week." The system drafts a note on the deal, a new contact, a commitment, a stage change and an email to Priya — each individually approvable. One utterance, five record updates, zero keystrokes.
Three groups spend most of their time here, and none of them are served by a phone:
The web also carries a commercial job. The buyer is usually the sales leader, and they evaluate on a laptop. A mobile-only demo can lose a deal to a competitor with a richer desktop surface even when our product is better for the people who use it daily. So the console must be credible and complete, not a settings page with a logo on it.
| Approach | Coverage | Trade-off |
|---|---|---|
| In-app dialer | Full recording and control | Rep must call from ORBIT — a habit change, but the prep brief gives them a reason to |
| Proxy-number bridging | Works with the native dialer | Extra step, caller-ID complexity, per-jurisdiction legality |
| Metadata plus voice memo | Universal fallback | No transcript; relies on a twenty-second note after the call |
Recommendation: in-app dialer as primary, built on Twilio Voice, with automatic post-call voice-memo prompting as the universal fallback. Meeting capture has no equivalent constraint and covers most high-value conversations anyway.
React Native with Expo for mobile: one codebase, over-the-air updates — critical when we are iterating on the queue weekly — and a hiring pool we can staff from. Native modules for telephony, push, background sync and secure storage. Offline-first, with the queue and entity cards cached locally and approvals queued optimistically. React and TypeScript on the web against the same API. Cold start to interactive Today under 1.5 seconds; approve-to-confirmed under 300 milliseconds perceived. Full screen-reader support and 44-point touch targets throughout.
Services, stores, latency budgets treated as objectives, and what must happen when each part fails.
| Path | p95 | Consequence if missed |
|---|---|---|
| Today render, cached | 400 ms | The three-minute day stops being true |
| Approve → confirmed | 300 ms | Perceived, via optimistic UI |
| Email received → visible | 30 s | Reps stop trusting the timeline |
| Call ends → ready follow-up | 90 s | Forces streaming transcription rather than batch |
| Ask anything → first token | 1.5 s | Feels broken past this |
| Report render from MV | 2 s | Managers go back to spreadsheets |
| Hourly spool completion | 10 min | The as-of stamp drifts from the hour boundary |
| Mode B sync propagation | 2 min | Credibility of the wedge |
Five hundred workspaces, roughly ten thousand seats. About two hundred activities per seat per day gives two million events a day — trivial for ClickHouse, comfortable for Postgres with time-based partitioning. Twenty hours of calls per seat per month is two hundred thousand hours a year of transcription. Fifty to two hundred thousand vectors per workspace puts us in the low tens of millions overall, within pgvector's range.
| Failure | Required behaviour |
|---|---|
| LLM provider down | Registry failover to the declared fallback; if all fail, serve the last computed queue with a staleness marker. Never a blank screen. |
| Bad prompt version shipped | Instant rollback to the pinned previous version. A registry row change, not a deploy. |
| Transcription backlog | Queue and retry; the meeting shows as processing; the ready follow-up arrives late with a notification |
| Mail provider rate limit | Backoff with jitter, and per-workspace fairness so one heavy tenant cannot starve others |
| Hourly spool fails | Serve the previous hour's views with the older as-of stamp visible, and alert. Reporting degrades in freshness, never in correctness. |
| ClickHouse unavailable | The transactional plane is unaffected — the rep's queue, records and drafts all keep working. Only reporting is down. |
| Mode B sync failure | Retry with backoff, surface in an admin health view, never drop the write — persist and replay |
Every AI behaviour here writes into a system a company runs on, or sends text to that company's customers. Quality infrastructure is the licence to operate.
If the system tells a rep "you promised them a 20 percent discount" and no such promise exists, and the rep acts on it with the customer, we have caused commercial damage to our customer.
The control: commitment extraction requires a verbatim transcript span, the span is stored, and the interface always exposes play the moment. A commitment without a playable source is never surfaced. A hard product invariant, not a guideline.
No customer data trains shared models without explicit contractual opt-in, and provider terms are verified for zero retention on all inference traffic. With enterprise buyers this is a differentiator, not a limitation.
Where the team's time goes, and what we deliberately refuse to spend it on.
Even fully staffed, engineering time is the scarcest resource in the company. The test for every line below is not "could we build this" — we could build all of it — but "does building it make the product better in a way a customer would notice, or does it just make us the owners of more code."
| Capability | Call | Reasoning |
|---|---|---|
| Provider framework | Build | The commercial wedge, and the thing that makes every later integration cheap. Nobody sells this shaped correctly for us. |
| Memory layer | Build | The differentiator. Off-the-shelf retrieval is precisely the naive design chapter 8 argues against. |
| Action engine and ready-state | Build | The product. This is what the user actually meets. |
| AI task registry | Build | Small, central, and the thing that keeps the system operable at scale. No vendor's abstraction fits our task taxonomy. |
| Salesforce and HubSpot sync | Build | Mode B is the wedge. Unified-API vendors are too shallow for bi-directional field-level conflict handling on the two targets that matter. |
| Google mail and calendar | Build | Highest-volume capture path, and the one where quota behaviour and reliability directly determine coverage. Worth owning. |
| Microsoft 365 mail and calendar | Buy first, build in parallel | Roughly half the enterprise market runs Microsoft, so coverage cannot wait for us to learn Graph. Behind a manifest entry, so the swap is a config change. |
| Meeting recording bots | Buy | Three vendor-specific bots maintained against continuously changing APIs, for zero differentiation. Abstract behind our own capture interface so we can bring it in-house if the unit economics ever justify it. |
| Transcription | Buy, multi-vendor | Commodity and improving fast. Register each vendor as a task provider so switching is a registry row. Benchmark on our own audio, including accented and code-switched English. |
| Enrichment | Buy | We own no dataset and should not build one. Priced into the margin rather than assumed away. |
| Telephony and dialer | Build on Twilio | The in-app dialer is how we capture calls at all on iOS, and recording control has to be ours. The carrier layer is bought; the product on top is not. |
| Long-tail CRM connectors | Buy | Low volume, not differentiating, not worth N native integrations |
| Auth, SSO, SCIM | Buy | Enterprise table stakes, zero differentiation, easy to get subtly and expensively wrong |
| Billing and subscriptions | Buy | Hosted checkout and metering. Do not build invoicing. |
| Vector storage | Build on pgvector | Our filtering needs are complex and our volume is moderate — pgvector's sweet spot. One tenancy model instead of two, with facts and embeddings written atomically. Revisit behind an interface if volume demands. |
| Semantic metric layer | Build | Small, core to reporting trust, and needs tight coupling to the AI query path |
| Predictive scoring models | Build | Gradient-boosted trees on our own features — cheaper, faster and explainable, which is what makes a sales leader act on a prediction |
| Observability, error tracking, status | Buy | Undifferentiated, and a safety system rather than a feature |
Own the capture path, rent the media plumbing. Coverage is tenet ten and it is where the product lives or dies, so the framework, the Google path and the CRM sync are ours. Bots, transcription and carrier minutes are interchangeable inputs, and owning them would buy nothing a customer can perceive.
Own anything the model reasons over. Memory, the registry, the semantic layer and the scoring models all shape what the system says to a customer. Renting any of those means renting our own quality ceiling.
Buying creates dependency, and dependency is a real risk when a vendor changes pricing, degrades, or is acquired. Four rules keep the exits open:
What the product in this document costs to build, how it sequences, and what the date does at different staffing levels.
This chapter does not change the scope to fit a team. It prices the scope, and shows the calendar consequence of each staffing choice, so the trade is made deliberately in this room rather than implicitly by whoever is available.
Estimates are in engineer-months for a competent engineer familiar with the domain, including tests, instrumentation and the operational work to run the thing — not just first-pass implementation. They assume no code carries over from anywhere.
| Workstream | Eng-months | Notes |
|---|---|---|
| Foundations — tenancy, auth, schema, RLS, CI scoping, environments | 6 | Unglamorous and load-bearing |
| Provider framework — manifest, webhook, poll, normaliser, health, gaps | 10 | Chapter 6. The multiplier on everything after it |
| Providers — Google, Microsoft, five funnel tools, two CRMs, long tail | 14 | Front-loaded; drops to ~0.5 each once the framework lands |
| Activity spine and identity resolution | 8 | Kafka, ordering, replay, merge ledger, adjudication |
| AI backbone — registry, routing, guardrails, evals, ready-state | 10 | Chapter 7 |
| Memory layer — L0 to L4, retrieval, citation gate, hygiene | 16 | Chapter 8. The largest single build, and the differentiator |
| Action engine — candidates, scoring, payload prep, feedback loop | 9 | Chapter 9 |
| Analytics plane — mirror, spool, semantic layer, MVs | 9 | Chapter 10 |
| Lead and pipeline modules, both modes | 12 | Mode A adds the full CRUD surface |
| Meeting module — capture, briefs, extraction, the 90-second loop | 10 | Includes the dialer |
| Workflow module — NL authoring, compiler, execution engine, templates | 12 | The deterministic engine is most of it |
| Reporting and analytics modules — Ask, proactive, predictive | 12 | Predictive needs calibration work, not just training |
| Native mobile — iOS and Android, offline, push, voice | 18 | Two platforms, plus the daily loop it exists for |
| Web application — full depth, admin, authoring, reporting | 14 | Not a console. A product |
| Design — two surfaces, a design system, and the loop itself | 14 | The differentiator is UX; this is not a support function |
| QA and eval operations — golden sets, sampling, labelling | 12 | Continuous, not a phase |
| Platform and SRE — deploys, observability, cost attribution, on-call | 10 | Grows with customers |
| Security and compliance — audit log, SSO, SCIM, SOC 2 readiness | 8 | Architectural, so it starts early |
| Total | ~204 | Plus roughly 15% integration and rework drag on a build this coupled |
Roughly 235 engineer-months with drag included. That is the price of the product described in this document — not an opening bid, and not padded.
| Function | Steady state | Note |
|---|---|---|
| Backend — platform and providers | 5 | Framework, spine, sync, identity. The largest group and the earliest hires |
| AI engineers | 3 | Registry, memory, extraction, retrieval, evals. Hardest to hire — start now |
| Data engineering | 2 | Analytics plane, semantic layer, spool, feature tables, predictive |
| Mobile | 3 | Two platforms plus the loop that carries the thesis |
| Web | 2 | The depth surface and much of the buying experience |
| Design | 2 | One product designer owning the loop, one owning breadth. Understaffing here produces a product indistinguishable from competitors |
| QA and eval operations | 2 | Human labelling and production sampling never stop |
| Platform and SRE | 1.5 | Kafka, ClickHouse, deploys, on-call |
| Engineering lead | 1 | Coding, not only coordinating, at this size |
| TPM | 1 | |
| Total | ~22 | Reached by month five, not on day one — see the ramp below |
The scope is identical across every row. Only the calendar and the parallelism change. Small teams lose less to coordination but cannot run workstreams concurrently; large teams run more at once but pay a widening coordination tax, and past roughly twenty-five people on a codebase this coupled, added engineers stop buying much time.
| Engineers | First sellable | Complete v1 | What it feels like |
|---|---|---|---|
| 2 | month 8 | month 34 | Competitors reach the market first. Viable only as a funded experiment, not as a category attempt |
| 6 | month 6 | month 18 | Serialised. One module at a time, mobile trails web by two quarters |
| 12 | month 5 | month 13 | Three concurrent workstreams. The first genuinely competitive row |
| 22 | month 4 | month 11 | Platform, modules, both surfaces and evals in parallel. The recommended row |
| 30 | month 4 | month 10 | Coordination cost eats most of the gain on a codebase this coupled |
Twenty-two people do not start on day one, and pretending otherwise would make the estimate fiction. The plan assumes two engineers in month one, eight by month three, sixteen by month five, and twenty-two by month seven — and the phasing below is arranged so the earliest work is the work a small group can do well while hiring runs in parallel.
Onboarding is priced in at roughly one engineer-month lost per hire. Hiring materially faster than this ramp does not pull the date in, because the codebase cannot absorb people faster than it can be explained.
Deliberately shaped for a small team, because this is when hiring is still running. Tenancy, schema, RLS and the CI scoping test; the provider framework end to end; the activity spine; identity resolution; the first four providers; and the timeline view a design partner can react to.
Exit: five partners connected across at least three different funnel tools; capture coverage above ninety-five percent; backfill under ten minutes; no silent gaps over a two-week observation.
The AI registry and ready-state; memory L0 to L3 with the citation gate; the action engine and the Today screen; meeting capture and the ninety-second post-call loop; mobile v1 on both platforms; Mode B two-way sync; the eval harness. This is the first sellable release.
Exit: daily action completion above fifty percent; call-end to ready follow-up under ninety seconds at p95; commitment extraction precision above ninety percent; first paying customer.
Lead and pipeline modules at full depth; Mode A as a system of record, with the CRUD surface, custom fields and pipeline configuration; the workflow module including natural-language authoring and the deterministic execution engine; the analytics plane and the reporting module; ten to fifteen providers.
Predictive scoring and forecasting; organisational memory; autonomy tier three for low-risk classes; coaching surfaces; SSO, SCIM, audit, residency and SOC 2 observation underway.
Per active seat per month, at twenty hours of calls, four hundred emails and forty meetings:
| Component | Estimate | Primary lever |
|---|---|---|
| Transcription | $3–5 | Skip calls under two minutes; batch pricing; multi-vendor via the registry |
| Model inference | $6–10 | Registry routing to cheap tiers; prompt caching on entity cards; batch API |
| Meeting bot infrastructure | $3–6 | Bought per hour; the one line we could insource later if volume justifies it |
| Enrichment | $2–5 | Bought at market — we own no dataset |
| Telephony | $1–3 | Carrier minutes on Twilio |
| Storage, databases, Kafka, ClickHouse | $3–5 | Two stores and a broker cost more than one database, and buy the isolation in chapter 10 |
| Total | $18–34 | At $99–149 per seat, roughly 70 to 80 percent gross margin |
Thinner than a platform that owns its enrichment data, and honestly so: we buy enrichment and meeting infrastructure rather than owning them. The price point sits higher to compensate, which is defensible for a product that removes a rep's administrative hour every day.
One. Route aggressively to cheap model tiers in the registry — most tasks are extraction and classification, not reasoning.
Two. Prompt-cache the entity cards, which sit in the prefix of nearly every prompt.
Three. Meter meeting hours, enrichment credits and call minutes into the plan allowance rather than absorbing unlimited usage. Per-seat base plus bundled credits with metered overage protects margin without punishing adoption.
Not a calendar. A dependency graph — what must exist before what, which things only look sequential, and the orderings that must never be inverted.
A schedule tells you when work happens. A build order tells you what breaks if it happens in the wrong sequence. The second is more durable, because it survives every change of date, and it is what lets several people work at once without discovering in month six that two of them built on top of a contract that was never agreed.
Contract before intake. Intake before meaning. Meaning before judgement. Judgement before surface. Everything else is parallel.
One chain runs the length of the project, and every week lost on it is a week lost overall:
tenancy → activity envelope → provider framework → spine → identity resolution → extraction → facts → entity cards → retrieval → action engine → Today screen → the post-call loop.
Nothing else is on it. Modules, reporting, Mode A, the workflow engine, the analytics plane and the second surface all hang off this chain and can be staffed in parallel the moment their upstream interface is stable. When staffing is scarce, it goes here first; when a decision is blocking, this is the chain to unblock.
These are the ones that cause teams to serialise work they could have run at the same time, and they are worth naming explicitly because the instinct is wrong in every case.
It is a schema, a resolver, a validator and a rollout mechanism. None of that requires a single captured activity. It can be built on day one alongside the provider framework, and it must be, because the moment the first extraction is written without it the prompts start scattering.
The mirror, the spool and the semantic layer depend on activity events existing — not on any module being finished. Data engineering can start the moment the envelope is agreed, rather than waiting for reporting requirements that will change anyway.
The CRUD surface, custom fields and pipeline configuration are ordinary application work depending only on the canonical entities. It is scheduled in phase two for commercial reasons, not technical ones, and can move earlier if a customer pulls it.
Tokens, the card anatomy, the gesture grammar and the empty-state discipline can be built and tested against fixtures. Design starting late is the most common way a product like this ends up looking like its competitors.
Each of these has a specific failure attached. They are cheap to honour and expensive to retrofit, which is the definition of an architectural decision.
| # | This before this | What happens if inverted |
|---|---|---|
| 1 | The activity envelope before the first provider | Provider number one's payload shape becomes the accidental schema, and every later provider is bent to fit a tool we happened to integrate first |
| 2 | Idempotency and dedupe keys before the second provider | Duplicate activities enter the spine, the memory layer double-counts, and the cleanup requires reprocessing history |
| 3 | The AI task registry before the first prompt | Prompts scatter into application code. Restoring the discipline later costs a full audit of every AI behaviour in the product |
| 4 | Identity resolution before extraction | Facts attach to the wrong person. Wrong facts are worse than missing facts, and unpicking them means re-deriving every card |
| 5 | The ready-state machine before the first outbound send | No audit trail of who cleared what, no uniform undo boundary, and no place to attach autonomy tiers later |
| 6 | Permission-aware retrieval before the second user in a workspace | A cross-user leak in a demo. Filtering after generation does not fix it — the model already saw the data |
| 7 | The eval harness before the second prompt version | Quality changes become unmeasurable, and every future regression is argued from anecdote |
| 8 | Deletion propagation before the first production customer | An erasure request arrives and we cannot honour it across embeddings, the mirror and downstream CRMs. This is a legal exposure, not a backlog item |
| 9 | The semantic layer before any AI querying of metrics | Text-to-SQL over raw tables hallucinates joins and returns two different answers to the same question. Trust does not recover from that |
| 10 | Coverage instrumentation before selling on the memory layer | We claim completeness we cannot measure, and the first confident wrong answer about a quiet account is discovered by a customer |
For each module: what it waits on, the order to build it in, and the thinnest slice that proves it works. The slice matters — it is what goes in front of a design partner before the module is finished.
Waits on: calendar provider, meeting-bot adapter, transcription task, identity resolution, extraction, entity cards, ready-state.
Order: calendar sync and meeting lifecycle → bot joins and recording lands in object storage → transcription with diarisation → speaker-to-contact mapping → summary → commitment extraction with source spans → drafted follow-up → the ninety-second notification loop → prep briefs.
First slice: one recorded meeting producing a summary and one correctly extracted commitment with a playable span.
Why first: it is the demo that sells the product, and it exercises every layer of the platform end to end. If meetings work, the platform works.
Waits on: provider framework, canonical entities, enrichment vendor, extraction.
Order: ingestion and dedupe → enrichment on arrival → lifecycle stages → reply intent classification → drafted first response → decay detection → assignment rules → learned ICP scoring last, because it needs closed-won history the workspace does not have on day one.
First slice: a lead arrives from a funnel provider, is enriched without human input, and produces one drafted response in the queue.
Waits on: canonical entities, CRM sync for Mode B, stage history, extraction.
Order: deal entity and immutable stage history → sync in from the customer's CRM → drafted field updates from conversation content → the cheap risk signals that need only the spine (no next step, silence, slipped dates, single-threading) → relationship mapping → forecasting last, because it depends on every signal above it being trustworthy.
First slice: a deal whose stage change was proposed from a call, with a citation, and approved in one tap.
Waits on: spine events, ready-state, action engine, registry.
Order: execution engine first, authoring second. Trigger taxonomy → deterministic evaluator → idempotent steps with per-entity locks → dry run → execution trace → blast-radius limits and kill switch → template library → natural-language authoring and the compiler → the match-count preview.
First slice: one hard-coded template — stalled deal produces a drafted check-in — running reliably with a visible trace.
The trap: building the authoring experience before the engine. Authoring is the demo; the engine is the product, and an authoring UI over an unreliable engine produces automations customers cannot trust.
Waits on: analytics plane, semantic layer.
Order: activity mirror → the hourly spool → metric definitions in code with tests → as-of stamping → the standard report library → scheduled delivery → drill-through to source records.
First slice: one metric — pipeline coverage — defined once, computed hourly, rendering identically in a report, a scheduled email and an Ask answer.
Waits on: reporting, memory retrieval, feature tables.
Order: Ask over the semantic layer for aggregate questions → Ask over memory for entity questions → the router that chooses between them → proactive rule-based insight → feature engineering during the spool → predictive models → calibration, published in-product → organisational memory last, because it needs volume rather than code.
First slice: one question — which deals are at risk this quarter and why — answered from the semantic layer and memory together, with citations.
Once the envelope contract is agreed and the provider framework has a stable interface, five tracks run concurrently and touch each other only through published contracts. This is the shape the team is staffed against in chapter 19.
| Track | Owns | Touches others only via |
|---|---|---|
| Intake | Provider framework, providers, spine, sync, identity | The activity envelope |
| Intelligence | Registry, extraction, memory, retrieval, evals | The envelope in, the entity card and fact API out |
| Judgement | Ready-state, action engine, ranking, feedback loop | The memory API in, the action API out |
| Surfaces | Design system, web, mobile | The public API and the design tokens |
| Data | Mirror, spool, semantic layer, feature tables, predictive | The spine topics and the metric definitions |
A block is not finished when it works. At this level of coupling it is finished when the next block can be built on top of it without archaeology:
Where the code lives, where the data lives, and the reasoning behind each boundary.
Three options, and the choice matters more than it looks, because the thing being shared across every codebase is the activity envelope and the API contract — and contract drift between five repositories is a slow, expensive class of bug.
| Option | Gains | Costs |
|---|---|---|
| One monorepo for everything | Atomic cross-cutting changes; one CI; no version skew | Heavy tooling investment early; mobile builds sit awkwardly inside it; CI times grow for everyone |
| A repo per service | Clean ownership, independent release cadence, simple CI | Contract drift; a schema change becomes a six-repo coordination exercise |
| Hybrid — server monorepo, published contracts, separate clients | Atomic changes where coupling is highest; independent release where it is lowest; one versioned source of truth for contracts | Requires discipline about what belongs in the contracts package |
Recommendation: a hybrid. One server-side monorepo where coupling is genuinely high, one contracts repository that every other repository depends on and none of them may bypass, and separate repositories for the clients and for anything with its own release cadence.
| Repository | Owns | Stack | Deploys to |
|---|---|---|---|
| orbit-contracts | The activity envelope, canonical entity schemas, Kafka event schemas, the OpenAPI spec, and generated clients for every language we use. The keystone. | JSON Schema, Protobuf or Avro, OpenAPI, codegen | Package registries |
| orbit-platform | Core API, domain model, provider framework and manifests, ingestion workers, identity resolution, sync fabric, ready-state, action engine, workflow execution engine, billing. A modular monolith with clear internal boundaries. | Ruby on Rails | App and worker services |
| orbit-ai | The task registry runtime, prompt sources, extraction, embedding, retrieval and ranking, scoring models, the eval runner. Exposes an internal HTTP API; owns no customer records. | Python, FastAPI | Inference and worker services |
| orbit-analytics | ClickHouse migrations, the mirror consumer, the hourly spool, metric definitions in code with their tests, feature tables for scoring. | SQL, dbt-style tooling, Python | Scheduled jobs, ClickHouse |
| orbit-web | The full web application — pipeline, reporting, provider setup and health, workflow authoring, admin, call review. | React, TypeScript | Static hosting or edge |
| orbit-mobile | The daily loop — Today, deals, Ask, voice, push, offline queue. | React Native, Expo | App stores plus over-the-air |
| orbit-design | Design tokens, the card anatomy, the gesture grammar, shared components for web and the primitives mobile mirrors. | TypeScript package | Package registry |
| orbit-infra | Infrastructure as code, environments, network policy, secrets wiring, alert routes, runbooks-as-code. | Terraform or Pulumi | Cloud accounts |
| orbit-evals | Golden datasets, labelling tooling, production sample review, quality dashboards. Access-restricted — see below. | Python, notebooks, datasets | Nothing; consumed by CI |
| orbit-docs | Architecture decision records, runbooks, on-call procedures, this specification and its successors. | Markdown | Internal site |
Golden datasets are built from real customer conversations. That makes them the most sensitive artifact we produce — more sensitive than the database, because they are portable, easy to copy into a notebook and easy to forget about.
They live in their own repository with restricted access, scrubbed of direct identifiers, covered by a retention policy, and referenced by CI rather than vendored into other repositories. Keeping them inside orbit-ai would put customer conversation text in front of every engineer who touches extraction.
Why provider manifests live inside orbit-platform rather than in their own repository. They are tightly coupled to the framework's execution semantics, and splitting them creates a version-matrix problem for no gain while we are the only authors. The important property is not repository separation but deployment separation: manifests and mappers load from a versioned artifact, so fixing a broken mapper does not require a full platform deploy. If third parties ever contribute providers, that is the moment to split.
Why prompts live in orbit-ai rather than in their own repository. A prompt without its output schema, its eval suite and its fallback chain is not a reviewable unit. Keeping them beside the registry runtime means a prompt change and its eval run in the same pull request. The active version is a registry row in the database; the repository is the source of truth and the review surface.
Why analytics is separate from platform. Different language, different release cadence, different failure tolerance, and a different owner. A metric definition change should not run the platform's test suite, and a platform deploy should not risk the reporting plane.
Six stores, each with a stated job and — more usefully — a stated non-job, because the failure mode here is data quietly ending up in the wrong place.
| Store | Holds | Explicitly does not hold |
|---|---|---|
| PostgreSQL primary | Tenancy and users; canonical entities; the activity table, partitioned by time; L2 facts with supersession; L3 entity cards; ready-state and audit; provider connections, cursors and secrets; workflow definitions and runs; the AI task registry; sync mappings and conflict log | Raw payloads and media; aggregate reporting queries; anything a request path scans in bulk |
| pgvector in the primary | Fact and card embeddings, in the same database so a fact and its vector are written in one transaction and share one tenancy model | A separate service. Revisit only when filtered recall degrades at volume |
| ClickHouse | The activity mirror; materialised views refreshed hourly; feature tables for scoring; anything that scans millions of rows across a few columns | The system of record. Nothing writes here first, and no request path reads from it live |
| Kafka | The activity spine, partitioned by workspace; replay for reprocessing when extraction improves; fan-out to memory, analytics, workflows and sync | Durable storage. Retention is a window, not an archive — the object store is the archive |
| Object storage S3-compatible | L0 raw artifacts — original payloads, mail bodies, recordings, transcripts, attachments, exports. Lifecycle-tiered, per-workspace encrypted | Anything queried directly. It is addressed by reference from Postgres, never scanned |
| Redis | Cache; rate-limit token buckets per connection and per provider; distributed locks for per-entity workflow concurrency; short-lived webhook dedupe; the application job queue | Anything whose loss matters. Everything here must be reconstructible |
Keyword search. Hybrid retrieval needs a lexical half, and vectors are weak on names, product terms and numbers. Start with Postgres full-text search rather than standing up OpenSearch — it is adequate at our early corpus size and it is one fewer stateful service. The trigger for moving is measurable: when lexical recall on entity names becomes the limiting factor in retrieval quality evaluations, not before.
The job queue. Kafka carries the spine; it should not carry application jobs. Retries, scheduling, priorities and visibility are different problems, and conflating them makes both harder to reason about. A Redis-backed queue for application work, Kafka for the event log, and a clear rule about which is which.
Every store above is regionally deployable, and the tenancy model — one workspace_id partition key across Postgres, pgvector and ClickHouse — is what makes a regional split tractable rather than a rewrite. If the geography decision in chapter 19 lands on a market with residency requirements, the shape is a full stack per region with a shared control plane, and it is far cheaper to design for now than to retrofit.
In dependency order, matching chapter 17: Postgres and object storage on day one; Redis with the first worker; Kafka when the second consumer appears, because a single consumer does not justify a broker; ClickHouse when the analytics track starts, which is as soon as the envelope is agreed. Nothing is stood up speculatively — each store arrives with the first workload that genuinely needs it.
What could kill this, what would make us stop, the ten decisions this room owns, and what we need.
| Risk | Severity | Mitigation |
|---|---|---|
| Understaffing without re-scoping — the plan runs at a lower row on the sensitivity table while the date stays fixed | Critical | Pick a row explicitly in this room. If staffing lands lower, the date moves or the scope is cut deliberately — never silently, and never by engineers absorbing it |
| Capture gaps make the memory lie — a provider stops delivering and nobody notices | Critical | Tenet ten, and the gap detection in chapter 6: cadence alarms, reconciliation counts, coverage as a user-visible number |
| AI trust collapse — one fabricated commitment in front of a customer's customer | Critical | Ready-state on everything; approval-first default; earned autonomy; the playable-span invariant |
| Hiring cannot keep the ramp — especially the three AI engineers | High | Start AI hiring before anything else. The phase-zero work is deliberately shaped for a small team so hiring runs in parallel rather than blocking |
| Well-funded competition reaches the market first | High | Sequence for a complete narrow promise at month four rather than a broad incomplete one at month eleven. Hold the four wedges rather than matching surface |
| Mobile-first meets a desktop buying process | High | A complete web application, not a console. Demo the post-call loop on a phone and the analytics on a laptop |
| Cost per seat runs away | Med-High | Per-workspace and per-task budgets from day one; cost per seat as a tracked objective; the three levers |
| Recording consent and data protection exposure | Med-High | Jurisdiction-aware consent shipped with recording; legal review before launch; deletion propagation across all stores designed in |
| Registry discipline erodes under deadline pressure | Medium | Lint rule plus review. Restoring it later costs a full audit of every AI behaviour in the product |
| Enterprise gates on a SOC 2 we do not have | Medium | Readiness work from phase one; sell mid-market first, where it is not a hard gate |
The single most consequential decision here. Scope is constant across rows; only the calendar moves. Recommend the twenty-two row — the first that puts a complete product in front of the market inside a year, and past which added headcount buys weeks rather than months.
Six modules at depth, both modes, native mobile and a full web application. If any of it is to be cut, it should be cut here and named, rather than discovered as a slip in month eight.
Recommend AI engineers and the product designer first, backend platform second, mobile third. The AI roles have the longest lead time and gate the largest workstream.
Recommend yes. Mode B is the easier first sale and generates the capture data; Mode A has to exist by the time the first customer asks to migrate, or the wedge is a dead end.
Recommend row-level with row-level security, one partition key across Postgres, pgvector and ClickHouse. Needs CTO sign-off.
Recommend buying for coverage while building the Google path natively, then reassessing. Half the enterprise market runs Microsoft and coverage cannot wait.
Recommend mid-market B2B teams of five to fifty reps, where there is no RevOps function forcing process compliance. Geography changes compliance, transcription and residency work materially, and blocks architecture.
Recommend per-seat base plus bundled AI, meeting, enrichment and call credits, with metered overage. Higher price point than a data-owning platform would need, for the reasons in chapter 19.
Recommend five, running at least three different funnel tools between them, so provider neutrality is proven rather than assumed.
Recommend the hybrid in chapter 18 — a server monorepo, a contracts repository nothing may bypass, separate clients — with Postgres, pgvector, ClickHouse, Kafka, object storage and Redis. Needs CTO sign-off before the first repository is created, because repository boundaries are the hardest thing on this list to change later.
Data model, competitive matrix, metric tree and the vocabulary this document uses precisely.
| Salesforce | HubSpot | Attio | Gong | Reevo | ORBIT | |
|---|---|---|---|---|---|---|
| AI-native architecture | no | no | partial | partial | yes | yes |
| Manual data entry | heavy | heavy | yes | n/a | minimal | near zero |
| Owns the record | yes | yes | yes | no | yes | Mode A |
| Works alongside an existing CRM | partial | no | no | yes | partial | Mode B |
| Provider-neutral intake | partial | partial | partial | no | own funnel | any tool |
| Native mobile for the daily loop | no | no | no | no | no | yes |
| Tells you what to do today | no | partial | no | partial | yes | core |
| Unified conversation memory | no | no | no | partial | yes | yes |
| Owned enrichment data | paid | paid | paid | no | yes | bought |
| Usable without training | no | partial | yes | yes | partial | the thesis |
North star — daily action completion rate: the share of surfaced items approved or marked ready the same day. Target 60 percent.
It simultaneously proves the items are correct (users trust them), useful (worth doing) and easy (low enough friction to finish). It cannot be gamed by surfacing more items, because it is a rate.
| Layer | Metric | Target |
|---|---|---|
| Capture | Interaction capture coverage | above 95% |
| Event to visible in timeline, p95 | under 30 s | |
| Distinct source types connected per workspace | 2 or more | |
| Memory | Fact extraction precision, labelled sample | above 90% |
| Answers with resolvable citations | 100% | |
| User correction rate on surfaced facts | under 5% | |
| Action | Daily action completion rate | above 60% |
| Acceptance rate by task type | above 40% | |
| Drafts approved without edit | above 50% | |
| Adoption | Rep mobile weekly active | above 70% |
| Manager web weekly active | above 80% | |
| Time to first value | under 10 min | |
| Data quality | Zero-input ratio | above 85% |
| Open deals with a scheduled next step | above 90% | |
| Forecast accuracy | beat the rep | |
| Economics | Cost per active seat | under $20 |
| Net revenue retention | above 110% | |
| Logo retention | above 90% |
| Term | Meaning in this document |
|---|---|
| Action | A ranked, prepared, expiring item in the daily queue, sitting in the drafted state |
| Activity spine | The single immutable normalised event log every provider writes to |
| Connection | One customer's authorised link to one provider, with its own cursor, secret, health and state |
| Dedupe key | Provider, connection and external identifier — what makes at-least-once delivery behave as effectively-once |
| Overlap window | The deliberate rewind behind a stored cursor that stops records falling into a gap |
| Provider | Any source of data, declared as a manifest rather than coded as an integration |
| Sweep | The low-frequency reconciliation poll that catches whatever the webhooks lost |
| AI task | A registered operation with a declared provider, model, prompt version, schema and eval suite |
| Analytics plane | The separate columnar store and its hourly-refreshed views; never queried by the application |
| Commitment | A promise extracted from a conversation, with a playable source span |
| Entity card · L3 | The rolling per-entity summary injected into every prompt and shown as the account view |
| Fact · L2 | An atomic, attributed, timestamped statement with a citation and a supersession chain |
| Mode A / Mode B | ORBIT as your CRM / ORBIT as an intelligence layer over the CRM you already run |
| Ready | The persisted state an item reaches when a human approves it, or edits it and marks it ready |
| Spool | The hourly job that moves deltas into the analytics plane and refreshes its views |
| Zero-input ratio | Share of record changes originating from capture rather than typing |
End of specification · v0.4 · Chapter 17 requires decisions before architecture is finalised