# Duplicate Requisition and Sourcing-Scan Playbook

This file pairs with the duplicate-guard stage of `job-seeker-resume-pipeline`. Use it whenever the user pastes a JD, recruiter DM, or job-board URL and the prior tailoring history is non-trivial.

## The problem this file solves

The "same requisition" can hide under three different disguises:

1. **Same URL, same week** — the user reposted the same UltiPro/Greenhouse/Requisitions link without remembering they pasted it days ago. Real example: Macmillan CAMPA003426 (2026-07-15) re-asked on 2026-07-23; the project already had a tailored triplet.
2. **Same role, different recruiter** — two recruiters representing the same client, each sending a paraphrased version. The job title and snippet differ, but the requisition is identical. Real example: Global Payments Marketing Solution Architect reposted by a second staffing firm on 2026-07-23; near-identical bullet text caught it.
3. **Same client, sister search** — Rob asks to scan Stripe's marketing/RevOps job board. The right answer is *not* to tailor 30+ resumes, it's to surface a ranked short list and tailor the top 1–3.

## Blocking first-step protocol (5 minutes, prevents wasted slots)

1. Extract the URL's stable identifiers: requisition/job ID, opportunity UUID, Greenhouse/UltiPro/LinkedIn IDs, and a distinctive 2–3 sentence phrase.
2. Run these in parallel before any preflight or generation:
   - `search_files` across `/root/Business_Projects/Project_1_Job_Seeker/`, `/root/.hermes/vault/01_Job_Seeker/`, `/root/.hermes/outbox/`, `/root/.hermes/staging/`. Search the company, the title, the recruiter, the requisition ID, and the distinctive phrase.
   - `session_search` for the same identifiers with `role_filter='user,assistant,tool'`, `sort='newest'`. The session DB is the most reliable place to find recruiter DMs and apply confirmations.
   - `read_file` on `/root/Business_Projects/Project_1_Job_Seeker/data/Rob_Blake_Job_Search_Tracker_V2.xlsx` and grep for the company+title.
3. **Textual-duplicate check.** Compare a 2–3 sentence snippet of the new JD to the snippet of any plausible prior JD. If the language is unusually close (same client intro, same bullet structure, near-identical phrasing), the new submission is almost certainly the same requisition. Surface the textual match to Rob before generating.
4. **Do not move on to preflight until Rob has confirmed whether the new outreach is the same role or a genuinely new one.** If it is the same, surface the existing vault memo, the existing DOCX, and the existing tracker status; ask whether Rob wants to re-submit, anchor on comp, or skip. If it is new, proceed with the standard pipeline.

## Textual-duplicate signals (2026-07-23 lessons)

- The two JDs share a client-specific intro line ("Our client is a leading...") or use the same team name ("Global SMB Marketing ecosystem," "SMB Marketing," "Treasurey/Treasury").
- The responsibilities section is structured identically and uses the same phrases ("ensure seamless data flow," "translate business and campaign requirements," "executive presence").
- Salary band is identical or near-identical, especially across different recruiters for the same client.
- The "qualifications" bullets list the same tools in the same order, even when the title is paraphrased.

When 3+ of these match, the JD is almost certainly a repost. Don't rely on title or recruiter alone.

## Recruiter-DM duplicate pattern

If the user pastes a recruiter DM (not a JD URL), the durable identifiers are weaker. The playbook:

1. Grep the recruiter name, the end client (if named), and any 2–3 sentence snippet in `session_search` first — past recruiter conversations live in the session DB, not the file tree.
2. If a prior conversation exists with the same recruiter for the same role, surface the comp-screening reply history and the rate already discussed. Don't draft a new opener.
3. If the recruiter is new but the role text is textually duplicate to a prior JD, treat the DM as a duplicate submission and ask Rob to confirm the end client and requisition ID before authorizing.

Real lesson (2026-07-23): the Ishan Ali recruiter DM was confirmed as a likely Global Payments Marketing Solution Architect repost only after the user asked for client confirmation. Had the textual-duplicate check been run at the duplicate-guard stage, that prompt would not have been necessary.

## Company-level sourcing scan (Stripe and similar)

When the user asks to scan a company's job board for roles that fit, the output is a single memo, not multiple tailored DOCXs.

### Workflow

1. **Extract the surfaced list** from the job-board page. The job-board search results are usually a single HTML table; `web_extract` returns the full table on a single call. Parse role, team, location, and job ID from the table.
2. **Dispatch 3 parallel `delegate_task` leaf agents** to inspect a curated subset of JDs in detail. Splitting the work in 3 parallel batches is the right tradeoff: fewer than 3 wastes time; more than 3 creates a coordination problem because each leaf returns its own summary.
3. **Synthesize the rankings** with: (a) role, (b) job ID, (c) base salary, (d) strongest verified matches, (e) hard gaps, (f) fit score 0–100, (g) apply/consider/skip verdict.
4. **Save the synthesis as a single vault memo** named `<Company>-US-Remote-Opportunity-Scan-<YYYY-MM-DD>.md`. This is the source of truth for the rest of the session and for future sessions. The user picks the top 1–3 from this memo.
5. **Run the standard tailoring pipeline only on the user's selections.** Do not pre-tailor everything "to be ready."

### Ranking heuristic (in order)

1. **Strongest verified evidence** — does the JD's required section have an honest, master-supported match? Stripe's "Stripe payment domain experience" is a *preferred* qualification, not a minimum, for the AI Accelerator role; that fact flipped the verdict from skip to apply. Always read the minimum section before treating preferred qualifications as hard gaps.
2. **Weakest honest hard gap** — for the strong matches, which hard requirement is the verified master *least* likely to support? That is the role's specific risk; surface it to Rob.
3. **Comp** — base salary and equity. Sort the strong matches by comp *within* the strongest evidence tier, not as the primary sort.
4. **Growth/scope** — is the role a one-off or a stepping stone? Strategist-level roles have more long-term leverage than IC niche roles for Rob's stated goal of moving toward the strategist level.

### Pitfalls specific to sourcing scans

- **Do not shotgun-tailor 5+ DOCXs from a single scan.** Each tailoring run is a non-trivial cost (cluster+phrase work, DOCX render, layout QA, vault memo, tracker entry). The right unit of work is a ranked short list, not a batch run.
- **Do not mix Stripe PMM, solutions architect, and MarOps roles in the same apply list.** Different resume templates, different ATS clusters, different hard requirements. The scan memo should group recommendations by resume template, not by raw rank.
- **Do not bypass the duplicate guard for sourcing scans.** The user's request to "scan a job board" can include roles already in the tracker. Run the duplicate guard against the tracker first.
- **The subagent summaries come back as truncated heads + tails.** Read `/root/.hermes/cache/delegation/subagent-summary-*.txt` directly when the synthesis needs the full middle, not just the head/tail.

## Worked example: Stripe search, 2026-07-23

The user asked: *"visit this webpage and scan for available jobs you think would fit me...knowing I'm going to lift the 'strategist' level focus to work for Stripe...https://stripe.com/jobs/search?teams=Customer+Success&teams=GTM+Partnerships&teams=Global+Strategic+Pursuits&teams=Marketing&teams=Revenue+Operations&teams=Sales&teams=Scaled+Sales&teams=Solutions+Architect&remote_locations=North+America--US+Remote"*

The correct flow:

1. **Surfaced the list** — `web_extract` on the search URL returned 41 US-remote roles across the 8 teams.
2. **Selected 12+ plausibly-fit roles** to inspect in detail. Roles like Account Executive, Account Development, and Recruiter were filtered out as obvious non-fits.
3. **Dispatched 3 parallel `delegate_task` leaf agents**:
   - Agent 1: marketing-operations, paid-digital, SEO/AI search, integrated campaigns.
   - Agent 2: PMM and lifecycle.
   - Agent 3: AI accelerator, RevOps, systems, payments strategist, solutions architect.
4. **Synthesized the rankings** with the heuristic above. All three agents independently ranked **Forward Deployed AI Accelerator, Marketing** as the standout.
5. **Saved the synthesis** to `Stripe-US-Remote-Opportunity-Scan-2026-07-23.md` in the vault.
6. **Recommended the top 1–3** instead of tailoring 5+ DOCXs: Forward Deployed AI Accelerator, Marketing → Campaign Operations Manager → Enterprise Paid Digital Marketing Manager.

The vault memo is reusable. If the user asks again in a future session to scan Stripe for strategist roles, read the memo first; only re-scan if it is stale (older than 2 weeks or the user reports new roles).
