# The AI SDR Is the Wrong Abstraction

Canonical URL: https://opengtm.co.uk/articles/the-ai-sdr-is-the-wrong-abstraction
Editor: [Oliver Longhurst](https://www.linkedin.com/in/oliver-longhurst)
Published: 2026-09-11
Updated: 2026-09-11
Standfirst: Treating prospecting as one autonomous rep hides the hard systems work. A better design is a composable pipeline whose durable advantage is the identity, evidence, policy and feedback it preserves.
Excerpt: The practical alternative to a monolithic AI SDR is a composable, evidence-bearing GTM control loop that can learn without locking the team to one model or vendor.

## LLM-readable summary

This article argues that an AI SDR is a misleading architecture for agentic prospecting because it conceals identity resolution, evidence provenance, policy and learning inside one synthetic role. It proposes a composable pipeline and a four-stage adoption path. The practical moat is the identity graph, evidence history, decision policy and outcome-labelled feedback loop, while models and tools remain replaceable.

Intended reader: GTM engineering, RevOps and revenue leaders evaluating agentic prospecting and mid-funnel systems
Practical implication: Instrument identity, evidence and outcomes first; add model recommendations next; automate only bounded action classes after their policy and evaluation gates are observable.

## Primary thesis

The AI SDR is the wrong abstraction because it bundles evidence gathering, identity, judgement, policy, action and learning into one opaque agent; a portable GTM control loop should separate those capabilities and make identity, provenance, permission and outcomes durable.

## Key claims

- A synthetic job title is a poor system boundary: it hides the interfaces that determine whether prospecting decisions are explainable, governable and learnable.
- A production GTM system should separate signal discovery, extraction, entity resolution, verification, decision policy, orchestration, engagement and measurement.
- The durable proprietary asset is the accumulated identity graph, evidence history, decision policy and outcome-labelled feedback rather than access to a particular model.
- Adoption should progress from observation to recommendation to bounded execution, with autonomy expanded by action class only after evidence supports it.
- Open repositories are most useful as replaceable capability boundaries; they do not by themselves establish production fitness, security, compliance or commercial outcomes.

## Caveats or limitations

- Repository capabilities were checked at immutable GitHub commits on 11 September 2026; this is a point-in-time architecture assessment, not an independent security audit or production benchmark.
- A public repository is not necessarily unrestricted open source. Review the exact licence, hosted-feature boundary, dependencies, data terms and trademark rules before commercial adoption.
- Entity matches, intent interpretations and model scores remain probabilistic. Consent, privacy, deliverability and sector-specific obligations require separate legal and operational review.
- The component set illustrates separable boundaries. It is not a recommendation to deploy or self-host every named project.

### LLM-summary caveats

- Repository capabilities were checked at immutable GitHub commits on 11 September 2026; this is a point-in-time architecture assessment, not an independent security audit or production benchmark.
- A public repository is not necessarily unrestricted open source. Review the exact licence, hosted-feature boundary, dependencies, data terms and trademark rules before commercial adoption.
- Entity matches, intent interpretations and model scores remain probabilistic. Consent, privacy, deliverability and sector-specific obligations require separate legal and operational review.
- The component set illustrates separable boundaries. It is not a recommendation to deploy or self-host every named project.

## Sources and citations

1. [changedetection.io README](https://github.com/dgtlmoon/changedetection.io/blob/07d00d081170f1d69aa1e11a91b164ed669ed2b8/README.md): changedetection.io contributors
1. [Crawl4AI README](https://github.com/unclecode/crawl4ai/blob/862f6bccb9c063f49b9d42701baa0eea17a4993f/README.md): Crawl4AI contributors
1. [Firecrawl README](https://github.com/firecrawl/firecrawl/blob/a54a9526b570ece4a6dd1adfa206a3462d4a030e/README.md): Firecrawl contributors
1. [Splink README](https://github.com/moj-analytical-services/splink/blob/2d57cd67e174a075c63db9904646c80c6613d727/README.md): UK Ministry of Justice Analytical Services
1. [Trigger.dev README](https://github.com/triggerdotdev/trigger.dev/blob/478422f4a38622cd0286c07a5e5f6a99825e1786/README.md): Trigger.dev contributors
1. [Nango README](https://github.com/NangoHQ/nango/blob/f2dc9d3aae60d8b5fefbd491cca9bbef1b02e688/README.md): Nango contributors
1. [Jitsu README](https://github.com/jitsucom/jitsu/blob/d88651c4ba8f04e9fcf6b950faed35ec622ad847/README.md): Jitsu contributors
1. [PostHog README](https://github.com/PostHog/posthog/blob/b9351db3f88e7fc181323e0053c8e285279c3ae8/README.md): PostHog contributors
1. [Twenty README](https://github.com/twentyhq/twenty/blob/123f463eea7501ed4b55a24b2312f46fd58638d0/README.md): Twenty contributors
1. [Papermark README](https://github.com/papermark/papermark/blob/ed19717ec02a1ac79aecf5569159aa9d2d869312/README.md): Papermark contributors
1. [Promptfoo README](https://github.com/promptfoo/promptfoo/blob/627bdd0cb70081be0c6c6583448dbf39127a0c70/README.md): Promptfoo contributors
1. [Docling README](https://github.com/docling-project/docling/blob/636c9d5c5ed0ba5707919817f97c44e18494c814/README.md): Docling Project contributors
1. [Stagehand README](https://github.com/browserbase/stagehand/blob/9f4f878e99ac82cd35480a7dd841dfe3dbb78093/README.md): Browserbase contributors
1. [Open Policy Agent README](https://github.com/open-policy-agent/opa/blob/8ce961d78fc6398ff25be63aa087d628c65927d2/README.md): Open Policy Agent contributors
1. [Presidio README](https://github.com/data-privacy-stack/presidio/blob/a7b17c75f3098b92b369f0b01855519f1cd5e8cc/README.MD): Data Privacy Stack contributors

## Concepts and context

Canonical summary: OpenGTM argues that an AI SDR is the wrong systems boundary for agentic prospecting. The durable design is a composable top- and mid-funnel pipeline that preserves an identity graph, evidence history, decision policy and outcome-labelled feedback while models and vendors remain replaceable.
Intended reader: GTM engineering, RevOps and revenue leaders evaluating agentic prospecting and mid-funnel systems
GTM stage: Top- and mid-funnel system design
Tags: GTM engineering, agentic systems, sales architecture, open-source software
Concepts: Composable GTM pipeline, Identity graph, Evidence history, Decision policy, Outcome-labelled feedback, Portable GTM Capability

## Body

The AI SDR is a seductive product category because everybody can picture the demo: give a synthetic rep a territory, a CRM login and an inbox, then ask it to research accounts, choose prospects and send messages. It is also the wrong systems abstraction.

A job title bundles judgement, context, policy, hand-offs and accountability into one human role. Turning that bundle into one agent hides the interfaces that determine whether the system learns or merely produces more activity. The result can look autonomous while remaining unable to explain why an account matched, which evidence was current, which policy allowed an action or whether the action improved a commercial outcome.

> The useful unit of design is a decision with evidence and an accountable outcome, not a synthetic employee.

## The pipeline is the product

A production GTM system is better understood as a sequence of replaceable capabilities. Each stage should accept typed inputs, preserve provenance, expose failure and hand a bounded result to the next stage.

1. Signal discovery: detect a material change or first-party event instead of repeatedly researching every account.

1. Extraction: retrieve the relevant page, document or application state and turn it into structured evidence.

1. Entity resolution: decide which company, person, domain and buying group the evidence belongs to, with confidence and merge history.

1. Verification and feature computation: distinguish observed facts from inferences, expire stale evidence and calculate decision inputs.

1. Decision policy: apply qualification criteria, consent, suppression, frequency limits, risk classes and human-review rules.

1. Orchestration and engagement: execute an approved next step with retries, idempotency and a visible system of record.

1. Measurement and feedback: connect the decision and action to a labelled outcome so prompts, rules and models can be evaluated later.

The language model may contribute at several points, but it is not the architecture. It is a replaceable reasoning component inside an evidence-bearing control loop.

## Why the monolith fails

### Bad identity becomes automated certainty

A CRM rarely starts with one clean representation of an account. Domains change, subsidiaries overlap, people use multiple addresses and product events arrive under incomplete identities. A monolithic agent tends to inherit those mismatches as facts. When research and outreach are coupled, one bad join can generate a confident but irrelevant message before anyone sees the mistake.

### The evidence disappears inside the answer

A prospecting brief without source, retrieval time and expiry is not durable account intelligence. It is a paragraph whose provenance will be forgotten. The system needs to retain what changed, where it was observed, when it was retrieved and which later decision consumed it.

### Policy becomes prompt folklore

Qualification rules, exclusions, consent, frequency limits and approval thresholds should be inspectable policy. If they live only in a long prompt, they are difficult to test, diff or enforce consistently across channels. The model may recommend; a policy boundary should decide what the system is permitted to do.

### Activity masquerades as learning

Replies, meetings and opportunities are not interchangeable labels, and a sent message is not a positive outcome. Without a decision-to-outcome record, the system cannot tell whether a score, evidence type, prompt or action improved revenue-generating conversations. It can only report that automation ran.

## The durable asset is the memory around the model

The proprietary advantage in this architecture is not access to a frontier model. Competitors can rent the same model, and model quality will keep moving. The harder-to-copy asset is the accumulated operating memory around it:

- **Identity graph.** Account, person, domain, product-user and buying-group relationships, including confidence, aliases and merge history.

- **Evidence history.** Time-stamped observations, source versions, extraction results, fact-versus-inference labels and freshness rules.

- **Decision policy.** Qualification logic, risk classes, consent and suppression rules, channel limits, escalation paths and human approvals.

- **Outcome-labelled feedback.** The trace from evidence to decision to action to outcome, including negative and ambiguous results rather than only wins.

Preserve those four things and you can replace a crawler, model, workflow engine, CRM or engagement tool without throwing away what the system has learned. Ignore them and every tool migration becomes an amnesia event.

## A curated component set, not a shopping list

The open repositories below are useful because they expose clean capability boundaries. They are examples to evaluate, not evidence that every team should self-host every layer.

### Signals and extraction

[changedetection.io](https://github.com/dgtlmoon/changedetection.io/blob/07d00d081170f1d69aa1e11a91b164ed669ed2b8/README.md) can turn monitored page changes into research triggers. Use either [Firecrawl](https://github.com/firecrawl/firecrawl/blob/a54a9526b570ece4a6dd1adfa206a3462d4a030e/README.md) for an API-led retrieval layer or [Crawl4AI](https://github.com/unclecode/crawl4ai/blob/862f6bccb9c063f49b9d42701baa0eea17a4993f/README.md) when Python integration and crawler control matter more. Start with one. Add [Stagehand](https://github.com/browserbase/stagehand/blob/9f4f878e99ac82cd35480a7dd841dfe3dbb78093/README.md) only for sources that genuinely require browser interaction, and [Docling](https://github.com/docling-project/docling/blob/636c9d5c5ed0ba5707919817f97c44e18494c814/README.md) when important evidence arrives as documents rather than ordinary web pages.

### Identity and transport

[Splink](https://github.com/moj-analytical-services/splink/blob/2d57cd67e174a075c63db9904646c80c6613d727/README.md) provides a serious probabilistic record-linkage primitive for entity resolution. [Jitsu](https://github.com/jitsucom/jitsu/blob/d88651c4ba8f04e9fcf6b950faed35ec622ad847/README.md) can collect first-party events, while [Nango](https://github.com/NangoHQ/nango/blob/f2dc9d3aae60d8b5fefbd491cca9bbef1b02e688/README.md) can isolate the recurring work of connecting to external applications. The useful design rule is to keep raw evidence and identity state outside the engagement tool that happens to consume them.

### Orchestration and the human-visible record

[Trigger.dev](https://github.com/triggerdotdev/trigger.dev/blob/478422f4a38622cd0286c07a5e5f6a99825e1786/README.md) is a code-first option for durable jobs, retries and scheduled work. [Twenty](https://github.com/twentyhq/twenty/blob/123f463eea7501ed4b55a24b2312f46fd58638d0/README.md) is an extensible CRM option when the team wants to control the account model and seller interface. An existing CRM can remain in place if it can display evidence and decisions without becoming the only copy of either.

### Behaviour and mid-funnel evidence

[PostHog](https://github.com/PostHog/posthog/blob/b9351db3f88e7fc181323e0053c8e285279c3ae8/README.md) can make first-party product and web behaviour inspectable. [Papermark](https://github.com/papermark/papermark/blob/ed19717ec02a1ac79aecf5569159aa9d2d869312/README.md) can add document and data-room engagement to the account timeline. Neither signal proves intent on its own; the value comes from combining observations with account context and recording how people interpreted them.

### Evaluation and policy

[Promptfoo](https://github.com/promptfoo/promptfoo/blob/627bdd0cb70081be0c6c6583448dbf39127a0c70/README.md) can turn prompt and model behaviour into repeatable evaluations. [Presidio](https://github.com/data-privacy-stack/presidio/blob/a7b17c75f3098b92b369f0b01855519f1cd5e8cc/README.MD) can support sensitive-data detection and transformation. [Open Policy Agent](https://github.com/open-policy-agent/opa/blob/8ce961d78fc6398ff25be63aa087d628c65927d2/README.md) is one option for separating declarative permission rules from model instructions. These controls still require a threat model, source-specific privacy decisions and human ownership; installing a repository is not governance.

## Adopt the system in four stages

### 1. Observe

Build the identity graph and evidence history before automating outreach. Capture a small set of high-value signals, show source-backed account briefs to humans and measure whether the evidence changes prioritisation. The first useful output is better judgement, not more messages.

### 2. Recommend

Let models propose account matches, qualification, next-best actions and draft language, but keep decisions visible and require human approval. Run new scores in shadow mode. Record overrides as data: disagreement is training signal, not workflow friction to be designed away.

### 3. Execute within bounds

Automate only reversible, low-risk actions whose policy can be tested: refresh research, create a task, route an account, prepare a draft or trigger an approved lifecycle step. Use idempotency, rate limits, suppression and explicit escalation. Expand autonomy by action class, not with one global switch.

### 4. Learn and promote

Join outcomes back to the evidence and decision that produced the action. Evaluate proposed prompt, policy and model changes against historical examples, including negative outcomes and human overrides. Promote a change only when it beats the current policy on agreed measures without crossing safety constraints.

## A practical architecture test

For any proposed GTM action, the system should be able to answer six questions:

- What changed or happened?

- Which source and source version support that observation?

- Why was the evidence linked to this account, person or buying group?

- Which decision rule or model version produced the recommendation?

- Which policy allowed the action, and was human approval required?

- What outcome followed, including no response, rejection or a human override?

If those answers cannot be reconstructed, the system is not an autonomous revenue capability. It is a prompt with a CRM login.

## Where humans should remain explicit

Human review is not a temporary embarrassment to remove after the demo. People should continue to own market definition, value propositions, consent posture, policy thresholds, ambiguous identity merges and high-impact external actions. The system can compress the cost of evidence gathering and make decisions more consistent; it cannot inherit accountability.

The boundary should move only when outcome data shows that a narrower action class is reliable and policy-compliant. That produces measured autonomy instead of theatre.

## Build the control loop, not the character

The AI SDR framing asks whether a model can imitate a representative from research to outreach. The better question is whether the GTM system can preserve identity, evidence, policy and outcomes while tools and models change underneath it.

Build that composable control loop and you can add autonomy where the evidence supports it. Buy the character first and the hardest parts remain hidden until they fail in public.
