# ATS & Audience Blueprint

**Purpose:** Step 1 knowledge base for a future dynamic resume/CV generator. This document defines audience mechanics, parsing behavior, labeling, matching, evidence controls, portfolio extraction, and zero-embarrassment guardrails. It does **not** generate a resume or CV.

**Version:** 1.1 — 13 August 2026

## 0. Core operating doctrine

There is no single universal “ATS score.” An ATS may parse a document, populate a candidate record, apply employer-configured knockout rules, support Boolean retrieval, calculate a skills-similarity score, integrate a third-party screener, or simply route the document to a human. These are different functions. For example, Workday documents a skills-match feature that weights required skills; Greenhouse documents both Boolean search and an optional assistive Talent Matching feature; Greenhouse also allows employer-configured application answers to trigger auto-rejection. Therefore, the generator must optimize for **accurate extraction, relevant retrieval, and rapid human comprehension**, never for an invented universal keyword-density formula. [Workday Candidate Skills Match](https://doc.workday.com/admin-guide/en-us/human-capital-management/recruiting/candidates/candidate-skills-match/bmj1604095304483.html), [Greenhouse Boolean Search](https://support.greenhouse.io/hc/en-us/articles/202360199-Search-candidates-using-Boolean-queries), [Greenhouse Talent Matching FAQ](https://support.greenhouse.io/hc/en-us/articles/41131616864283-Talent-Matching-Data-Processing-FAQ), [Greenhouse Auto-Reject](https://support.greenhouse.io/hc/en-us/articles/360000653472-Auto-reject).

| ID | Insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| D-01 | Truth outranks optimization | **Do:** Preserve verified titles, dates, technologies, responsibilities, and outcomes. **Do not:** add a keyword, metric, seniority level, or claim merely to improve apparent match. | **Do:** Preserve exact authorship, publication status, methods, findings, funding status, and teaching role. **Do not:** upgrade “submitted” to “accepted,” “assistant” to “instructor,” or a proposal to a completed finding. |
| D-02 | Explicit instructions outrank defaults | **Do:** Follow the vacancy’s requested file type, page limit, fields, and naming convention before these defaults. **Do not:** force PDF, two pages, or a standard section when the employer explicitly says otherwise. | **Do:** Follow the program, lab, institution, or funder template exactly. **Do not:** submit a general academic CV where NIH, NSF, ERC, UKRI, or a portal requires a specific biosketch/narrative format. |
| D-03 | Two document models must remain separate | **Do:** Generate a selective, role-specific 1–2 page resume. **Do not:** dump the candidate’s complete academic history into an industry application. | **Do:** Generate a comprehensive, discipline-aware CV whose ordering is tailored to the opportunity. **Do not:** compress a research career into corporate accomplishment bullets or assume one page is superior. |
| D-04 | Defaults are not universal laws | **Do:** Treat these rules as safe defaults across Workday, Greenhouse, Lever, Taleo, iCIMS, and common parser integrations. **Do not:** claim guaranteed passage through every tenant, licensed module, or custom screening tool. | **Do:** Treat committee behavior as institution-, discipline-, programme-, and funder-dependent. **Do not:** claim a universal academic scoring algorithm or a guaranteed admissions/funding outcome. |
| D-05 | Submission-ready means no unresolved uncertainty | **Do:** Omit unsupported claims or ask the user before finalization. **Do not:** expose `[VERIFY]`, invented placeholders, or silent guesses in a final resume. | **Do:** Resolve publication status, author order, dates, supervisors, award status, and reference consent before finalization. **Do not:** leave ambiguous scholarly claims in a submitted CV. |

## 1. Audience and decision dynamics

Corporate documents move through configurable software and several human roles. Workday describes separate review, screen, assessment, and interview subprocesses; Greenhouse scorecards evaluate predetermined skills, traits, qualifications, and other attributes. Academic applications are read as a dossier: CV, statement, transcript, letters, and sometimes writing samples or publications. MIT and Cornell guidance emphasize research preparedness, concrete experience, future direction, and fit with the program or lab rather than raw term frequency. [Workday recruiting subprocesses](https://doc.workday.com/workday-education/en-us/course-manuals/recruiting-for-administrators/recruiting-subprocesses.html), [Greenhouse scorecards](https://support.greenhouse.io/hc/en-us/articles/4414777492891-Scorecard-overview), [MIT graduate statement guidance](https://mitcommlab.mit.edu/eecs/commkit/graduate-school-statement-of-purpose/), [Cornell academic statement guidance](https://gradschool.cornell.edu/inclusion/recruitment/prospective-students/writing-your-statement-of-purpose/).

| ID | Insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| A-01 | Primary document purpose | **Do:** Make relevant capability, production contribution, and business/engineering impact obvious quickly. **Do not:** treat the resume as a complete biography. | **Do:** Document scholarly preparation, trajectory, outputs, methods, teaching, recognition, and service. **Do not:** treat the CV as a sales flyer. |
| A-02 | First software layer | **Do:** Make the file text-readable and the data structurally mappable. **Do not:** assume attractive visual rendering means correct machine extraction. | **Do:** Complete every structured application field separately and accurately. **Do not:** assume the uploaded CV will populate or override portal fields. |
| A-03 | Eligibility/knockout layer | **Do:** Answer work authorization, location, license, compensation, availability, and similar application questions truthfully; they may be configured as automatic gates. **Do not:** try to compensate for an ineligible answer through résumé wording. | **Do:** Meet degree, prerequisite, language, deadline, supervisor, and funding eligibility rules explicitly. **Do not:** assume research fit can waive administrative ineligibility. |
| A-04 | HR/generalist scan | **Do:** Surface recognizable role identity, years/scope where verified, core stack, domain, eligibility, and relevant outcomes in plain language. **Do not:** lead with an unexplained internal project name or niche acronym. | **Do:** Make current academic stage, field, research direction, institution, and major outputs easy for a broad committee member to identify. **Do not:** open with a dense specialist taxonomy that only one PI can decode. |
| A-05 | Engineering-manager scan | **Do:** Show ownership boundaries, architecture decisions, scale, reliability, security, delivery, testing, and operational evidence. **Do not:** offer a tools inventory with no evidence of how the tools were used. | **Do:** Show the research question, method, individual contribution, evaluation, finding/status, and limitations. **Do not:** replace scholarly reasoning with a production-stack list. |
| A-06 | Human risk detection | **Do:** Reduce uncertainty about what was built, the candidate’s role, and whether impact is credible. **Do not:** use inflated verbs that force the reader to question ownership. | **Do:** Reduce uncertainty about authorship, independence, methodological competence, and research continuity. **Do not:** blur team output with the applicant’s personal contribution. |
| A-07 | Corporate evidence hierarchy | **Do:** Prioritize recent role-relevant production work, verified outcomes, and technical decisions; keep education/certifications proportionate. **Do not:** let older or weaker material displace stronger evidence. | **Do:** Prioritize education, research experience, publications/outputs, teaching, and awards according to the opportunity and discipline. **Do not:** default to employment chronology when research credentials are the selection basis. |
| A-08 | Academic fit psychology | **Do:** If applying to industrial research, translate research into the employer’s problem, methods, and deliverables. **Do not:** name-drop a lab or paper without a concrete connection. | **Do:** Demonstrate fit as a chain: target question → relevant method/domain → prior evidence → credible next research step → lab/program resources. **Do not:** equate repeated faculty names or topic words with genuine alignment. |
| A-09 | References and letters | **Do:** Omit references and “References available upon request” unless the employer asks. **Do not:** consume scarce resume space with referee details. | **Do:** Treat informed academic references/letters as independent high-value evidence and include referee details only when requested or customary, with consent. **Do not:** fabricate contact details or list someone who has not agreed. |
| A-10 | Length | **Do:** Target one page for early-career candidates and up to two focused pages when relevant experience justifies it. **Do not:** shrink type or margins merely to satisfy a page count. | **Do:** Allow the CV to reach the length needed for a clear scholarly record; two to several pages is normal. **Do not:** interpret “comprehensive” as permission for repetition, irrelevant detail, or unreadable density. |
| A-11 | Funding-track exception | **Do:** For an R&D employer, still use the requested corporate format unless a research CV is requested. **Do not:** assume “research” automatically means academic CV. | **Do:** Use the exact funder format. NIH and NSF now use standardized common-form biosketches; ERC and UKRI may require bounded or narrative track-record formats. **Do not:** submit an unrestricted multi-page CV when a scheme-specific document is required. |
| A-12 | Time-to-understanding | **Do:** Design for fast triage without citing an unverified universal “six-second rule.” **Do not:** build policy around a marketing statistic. | **Do:** Put the strongest and most relevant scholarly signals early, while preserving complete records later. **Do not:** assume a professor reads every line before forming an initial view. |

## 2. Parsing architecture and the professor-PDF contrast

Modern parsing is best modeled as a pipeline, although vendor implementations differ and much of their internal architecture is proprietary. Textkernel publicly describes conversion of the source document to plain text, usability checks, optional OCR, parsing, normalization, and structured JSON output. Its data model can compute information such as total experience and highest education; its taxonomies normalize skills and professions. DaXtra reports extraction of 150+ fields into XML or JSON with skills taxonomies. Greenhouse’s current Talent Matching uses narrowly scoped models to extract skills, titles, years of experience, dates, and employers. [Textkernel parsing workflow](https://developer.textkernel.com/tx-platform/v10/resume-parser/overview/getting-started/), [Textkernel data model](https://developer.textkernel.com/Parser/master/data_model/), [DaXtra Parser](https://www.daxtra.com/products/resume-parsing-software/), [Greenhouse Talent Matching FAQ](https://support.greenhouse.io/hc/en-us/articles/41131616864283-Talent-Matching-Data-Processing-FAQ).

| ID | Pipeline stage / insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| P-01 | File ingestion | **Do:** Supply the original, uncorrupted DOCX or text-based PDF in the requested format. **Do not:** upload a screenshot, scanned printout, recruiter-marked copy, or repeatedly converted file. | **Do:** Export a stable, searchable PDF from the source document unless the portal requests another format. **Do not:** upload image-only pages or a document whose fonts/symbols render inconsistently. |
| P-02 | Text conversion and OCR | **Do:** Ensure selectable text and test extraction before submission. **Do not:** rely on OCR to rescue decorative or scanned content; conversion failure occurs before semantic parsing. | **Do:** Ensure names, paper titles, formulas, symbols, and citations remain searchable and readable. **Do not:** assume a human reviewer will manually repair broken text or missing glyphs. |
| P-03 | Reading-order reconstruction | **Do:** Use a single linear content flow. **Do not:** use columns, floating text boxes, sidebars, or overlapping elements that can interleave lines. | **Do:** Prefer a clean one-column scholarly layout even though humans tolerate more design variation. **Do not:** make the reader jump between unrelated columns or hunt for dates and statuses. |
| P-04 | Section segmentation | **Do:** Use canonical headings so experience, education, skills, and projects have reliable boundaries. **Do not:** replace headings with creative labels such as “My Journey,” “Toolbox,” or “Where I’ve Made Magic.” | **Do:** Use discipline-recognized categories such as Research Experience, Publications, Teaching Experience, and Grants and Fellowships. **Do not:** merge distinct scholarly statuses under vague headings. |
| P-05 | Token/entity extraction | **Do:** Spell out the complete entity and add a standard acronym where useful: “Amazon Web Services (AWS).” **Do not:** split entity names with decorative spacing or symbols. | **Do:** give exact institution, degree, venue, award, method, and supervisor names. **Do not:** abbreviate niche venues or methods before defining them. |
| P-06 | Relationship linking | **Do:** Keep employer, title, location, dates, and bullets in one contiguous entry. **Do not:** place all employers in one column and all dates/titles in another. | **Do:** keep each research role, institution/lab, supervisor, dates, methods, and contribution together. **Do not:** leave it unclear which supervisor, project, or output belongs to which role. |
| P-07 | Temporal interpretation | **Do:** use explicit `Month YYYY – Month YYYY` or `Month YYYY – Present` consistently. **Do not:** infer “Present” from a missing end date or use ambiguous numeric dates. | **Do:** use a consistent year or month-year convention appropriate to the record and identify expected completion explicitly. **Do not:** make future, ongoing, and completed work indistinguishable. |
| P-08 | Taxonomy normalization | **Do:** use truthful canonical names plus common synonyms only where they aid retrieval, such as “continuous integration and continuous delivery (CI/CD).” **Do not:** add adjacent skills because a taxonomy might associate them. | **Do:** use standard discipline terms, method names, and classifications accurately. **Do not:** broaden a narrow method into a field-level capability the applicant has not demonstrated. |
| P-09 | Structured serialization | **Do:** assume extracted data will become arrays of employment, education, skills, certifications, links, and other fields. **Do not:** depend on visual proximity alone when labels are ambiguous. | **Do:** assume the committee may compare the PDF with separately structured education, publication, and referee fields. **Do not:** tolerate contradictions between the CV and portal record. |
| P-10 | Indexing and retrieval | **Do:** include relevant exact terms naturally in Technical Skills and in evidence-bearing Experience/Projects. **Do not:** hide terms, repeat blocks, or stuff white text. | **Do:** make research topics and methods easy to locate through clear headings and precise wording. **Do not:** write for a hypothetical academic keyword score. |
| P-11 | Matching/scoring | **Do:** distinguish parse success from match quality. A document can parse perfectly and still lack required evidence. **Do not:** describe the ATS as a single robot that “accepts” or “rejects” every resume. | **Do:** distinguish administrative screening, committee judgment, PI interest, funding capacity, and final selection. **Do not:** claim the CV alone determines admission or funding. |
| P-12 | Human verification | **Do:** review auto-filled fields after upload whenever the portal permits; Workday and Greenhouse both warn parsing can vary or fail. **Do not:** submit without checking name, contact, employer, title, and dates. | **Do:** preview the final uploaded PDF and every form field in the portal. **Do not:** assume the locally rendered file and the portal preview are identical. |

### 2.1 2026 vendor reality matrix: Workday, Greenhouse, Lever, Taleo, and iCIMS

This matrix records only behavior the vendors currently document publicly. It does not imply that every customer enables every module. A product can parse, retrieve, rank, gate, and investigate fraud through separate components, and an employer may add third-party tools or custom workflow rules.

| ID | Vendor / documented mechanics | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| V-01 | **Workday:** resume parsing can populate candidate fields; Candidate Skills Match is a separate feature that weights required skills. Workday's own June 2026 recruitment notice also describes extracting skills, experience, location, and education and assigning suggested qualification grades in its internal hiring process—evidence that matching can support prioritization, not proof that every Workday customer uses the same configuration. [Parsing](https://doc.workday.com/admin-guide/en-us/human-capital-management/recruiting/candidates/set-up-prospects-and-candidates/hdc1552497830785.html), [Skills Match](https://doc.workday.com/admin-guide/en-us/human-capital-management/recruiting/candidates/candidate-skills-match/bmj1604095304483.html), [Workday recruitment notice](https://www.workday.com/en-us/privacy/recruiting-privacy-statement.html). | **Do:** optimize clean field extraction and truthful evidence for required qualifications. **Do not:** present Workday's own hiring configuration as a universal Workday tenant algorithm or reverse-engineer an invented threshold. | **Do:** treat a Workday-hosted university application as a structured portal plus uploaded dossier. **Do not:** assume a professor sees, trusts, or uses a Workday match grade. |
| V-02 | **Greenhouse:** parser failure guidance identifies specific layout risks; Boolean search, employer-configured auto-reject questions, and optional Talent Matching are distinct. The June 2026 Talent Matching documentation says employers define and weight calibration criteria, match output is assistive, and it does not automatically advance or reject candidates. [Parsing failures](https://support.greenhouse.io/hc/en-us/articles/200989175-Unsuccessful-resume-parse), [Boolean search](https://support.greenhouse.io/hc/en-us/articles/202360199-Search-candidates-using-Boolean-queries), [Auto-reject](https://support.greenhouse.io/hc/en-us/articles/360000653472-Auto-reject), [Talent Matching](https://support.greenhouse.io/hc/en-us/articles/41396009937307-Talent-Matching). | **Do:** satisfy truthful application answers, parser-safe structure, and evidence-bearing criteria separately. **Do not:** confuse an unsuccessful parse, an application-question gate, a search result, and an AI match category. | **Do:** complete structured fields and preserve a readable PDF if an academic employer uses Greenhouse. **Do not:** optimize an admissions CV for Greenhouse's optional corporate Talent Matching semantics unless the institution explicitly says it uses them. |
| V-03 | **Lever:** public documentation describes resume parsing into candidate profiles and database retrieval; Lever's current product material also advertises AI-assisted candidate ranking and fraud signals. Exact production logic and customer configuration remain proprietary. [Resume parsing](https://help.lever.co/s/article/Understanding-Resume-Parsing), [Candidate search](https://help.lever.co/s/article/Searching-the-Database-for-Candidates), [Current platform](https://www.lever.co/why-lever). | **Do:** use explicit titles, skills, employers, and dates that can populate and retrieve as fields. **Do not:** claim a particular Lever score, cut-off, fraud outcome, or guaranteed ranking from marketing-level feature descriptions. | **Do:** assume a Lever-hosted academic hiring application can expose the CV to recruiting software before human review. **Do not:** collapse faculty assessment into Lever retrieval or ranking. |
| V-04 | **Oracle Taleo:** Oracle's publicly accessible Taleo Enterprise documentation describes third-party resume parsing; extraction of personal, education, and work fields; conceptual search over resume/job-description text; and separate application-flow screening criteria. The accessible manuals are legacy 20x/21x documentation, not evidence of a newly released 2026 engine, and limits/settings may be tenant-configurable. [Candidate Management](https://docs.oracle.com/en/cloud/saas/taleo-enterprise/21b/otrec/candidate-management.html), [Application Flow Blocks](https://docs.oracle.com/en/cloud/saas/taleo-enterprise/otcug/r-applicationflowblocks.html), [Getting Started](https://docs.oracle.com/en/cloud/saas/taleo-enterprise/20d/otrec/getting-started.html). | **Do:** preserve standard field boundaries and natural conceptual relevance. **Do not:** quote a legacy file-size default, supported-language list, or parser behavior as a universal current Taleo rule; do not confuse conceptual search with knockout screening. | **Do:** follow the institution's live Taleo portal instructions and verify parsed fields. **Do not:** infer how an academic committee evaluates the PDF from legacy Taleo search documentation. |
| V-05 | **iCIMS:** its June 2026 ATS guide documents application parsing that makes resumes/CVs searchable and can prepopulate fields, plus keyword/natural-language search and AI screening/matching. Current product pages describe candidate comparison, ranking, matching, and configurable workflows. [2026 ATS guide](https://www.icims.com/blog/what-to-look-for-in-an-applicant-tracking-system-complete-buyers-guide/), [AI recruiting](https://www.icims.com/products/ai-recruiting-software/), [Enterprise ATS](https://www.icims.com/products/hiring-software/enterprise-applicant-tracking-system/). | **Do:** make required skills and experience unambiguous in both the resume and application fields. **Do not:** assume every iCIMS customer licenses the same modules, uses the same ranking, or weighs the same criteria. | **Do:** treat iCIMS parsing as an administrative/retrieval layer when used by a university. **Do not:** substitute corporate keyword tactics for research-method, output, and lab-fit evidence. |
| V-06 | Cross-vendor safest common denominator | **Do:** build one linear, text-based, single-column master and validate the portal's parsed result. **Do not:** optimize for a vendor-specific quirk that degrades truth, readability, or another vendor's parsing. | **Do:** preserve a human-first scholarly PDF while keeping text/searchability and portal fields consistent. **Do not:** assume the safest corporate layout requires stripping useful academic detail or discipline-standard citations. |

## 3. Data fields and mapping rules

Parsers do more than search text; they classify text into schemas. Textkernel documents normalized fields and derived attributes, while DaXtra describes contact, work, education, and skill extraction. The frequently cited Sovren parser is now part of Textkernel following its acquisition and brand consolidation, so current Textkernel/Tx Platform documentation is the relevant continuation for the Sovren-style schema and matching model. The generator should therefore format each entry as a self-contained record. [Textkernel–Sovren acquisition](https://www.textkernel.com/learn-support/blog/textkernel-acquires-sovren-to-become-the-global-leader-in-ai-powered-recruitment-technology/).

| ID | Field / mapping risk | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| M-01 | Identity and contact | **Do:** put full name, city/country as appropriate, professional email, phone, LinkedIn, and portfolio/GitHub in the main top body. **Do not:** place essential contact data only in a header, footer, image, or icon. | **Do:** provide professional contact data and relevant identifiers such as ORCID, Google Scholar, or a research site when useful. **Do not:** include marital status, age, photograph, or unrelated sensitive data unless a jurisdiction explicitly requires it. |
| M-02 | Employment record | **Do:** use `Employer — Role — Location — Dates`, followed by contribution bullets. **Do not:** disguise the legal employer as a client or silently rewrite the official title. | **Do:** distinguish industry employment, research appointments, visiting roles, and assistantships accurately. **Do not:** collapse appointments with materially different scholarly status. |
| M-03 | Education record | **Do:** state institution, exact degree, field, location if useful, and completion/expected date; add GPA only when favorable, relevant, and accurate. **Do not:** imply a completed degree when it is in progress. | **Do:** include degree, field, institution, dates, thesis/dissertation title, supervisor, and distinctions where relevant. **Do not:** invent a degree translation or equivalency. |
| M-04 | Skill record | **Do:** group skills by type—Languages, Frameworks, Data, Cloud/DevOps, Testing/Observability—and support priority skills in experience. **Do not:** use rating bars, stars, percentages, or unsupported proficiency labels. | **Do:** distinguish research methods, programming/tools, laboratory techniques, and human languages. **Do not:** let a generic skill cloud substitute for evidence in Research Experience. |
| M-05 | Project record | **Do:** state project context, individual contribution, relevant technology, scale/constraint, validation, and outcome. **Do not:** list a project name and stack with no action or evidence. | **Do:** state research problem, method, dataset/system, individual contribution, result/status, supervisor/collaborators, and output. **Do not:** report a hypothesis as a result. |
| M-06 | Publication/output record | **Do:** include only selected role-relevant publications, patents, talks, or open-source outputs and label them accurately. **Do not:** crowd a software resume with an exhaustive bibliography. | **Do:** use discipline-standard citations and separate peer-reviewed publications, preprints, manuscripts under review, talks/posters, software, datasets, and works in progress as appropriate. **Do not:** mix statuses to inflate the publication record. |
| M-07 | Certification/award record | **Do:** include exact credential, issuer, date, and expiry/ID when relevant and safe. **Do not:** list training attendance as certification. | **Do:** include exact award/fellowship/grant name, funder, role, status, year, and verified amount when appropriate. **Do not:** call a nomination, application, or pending decision an award. |
| M-08 | Link record | **Do:** display a short readable URL or recognizable domain and verify the destination. **Do not:** rely on anchor text whose hidden target may be lost in parsing. | **Do:** link DOI, publication, dataset, software, ORCID, or lab profile where it adds verification. **Do not:** link inaccessible drafts, broken pages, or confidential repositories. |
| M-09 | Location and work mode | **Do:** state location/remote status only when it clarifies eligibility or employment context. **Do not:** infer relocation willingness or work authorization. | **Do:** state institutional location and research visit context accurately. **Do not:** infer citizenship, visa status, or geographic eligibility from education/employment. |
| M-10 | Derived fields | **Do:** allow calculations only from complete, non-overlapping source data and preserve units/time windows. **Do not:** infer total experience by summing overlapping roles or infer mastery from elapsed time. | **Do:** derive counts or durations only when the underlying record is complete and the calculation matters. **Do not:** infer citation impact, independence, or research quality from publication count alone. |

## 4. Formatting and document-parsing guardrails

Formatting failures are not folklore. Greenhouse explicitly lists images, word art, complex tables, headers/footers, text boxes, columns, unclear sections, and incomplete titles among causes of failed or partial parsing. Workday advises avoiding images/image-based styles, and Penn recommends a simple single-column design for ATS readability. Academic committees tolerate more visual variation, but university guidance still favors simple, readable, professional CVs. [Greenhouse parsing failures](https://support.greenhouse.io/hc/en-us/articles/200989175-Unsuccessful-resume-parse), [Workday Resume Parsing](https://doc.workday.com/admin-guide/en-us/human-capital-management/recruiting/candidates/set-up-prospects-and-candidates/hdc1552497830785.html), [Penn application guidance](https://careerservices.upenn.edu/internship-and-job-applications-for-masters-students/), [Penn faculty CV guidance](https://careerservices.upenn.edu/application-materials-for-the-faculty-job-search/cvs-for-faculty-job-applications/).

| ID | Formatting insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| F-01 | File type | **Do:** use the employer-requested format; otherwise use a text-based PDF or clean DOCX supported by the portal. **Do not:** claim PDF is always superior—some workflows explicitly request DOCX. | **Do:** normally use a searchable PDF to preserve scholarly layout, unless instructions specify DOCX or a generated form. **Do not:** disregard a required template. |
| F-02 | Page geometry | **Do:** use standard page size for the target country, margins roughly 0.65–1 inch, and enough white space. **Do not:** compress margins until extraction or reading becomes difficult. | **Do:** use stable margins around 0.75–1 inch and consistent spacing across pages. **Do not:** create visually exhausting walls of citations or entries. |
| F-03 | Layout | **Do:** use one column from top to bottom. **Do not:** use a sidebar, two-column template, infographic résumé, or floating blocks. | **Do:** prefer a straightforward one-column CV; aligned dates are acceptable when reading order remains obvious. **Do not:** use magazine-style composition or ornamental sidebars. |
| F-04 | Typography | **Do:** use a professional, widely available font such as Arial, Calibri, Helvetica, Georgia, or Times New Roman at a readable 10–12 pt body size. **Do not:** use novelty fonts, ultra-light weights, or tiny type. | **Do:** use a readable professional font, normally 10.5–12 pt; discipline/institution conventions may govern. **Do not:** mix many font families or use typography as decoration. |
| F-05 | Font encoding | **Do:** embed fonts in PDF and verify ligatures/symbols survive text extraction. **Do not:** use fonts that turn letters into shapes or cause missing glyphs. | **Do:** verify mathematical symbols, non-Latin names, accents, and bibliography characters after export. **Do not:** submit a PDF with substituted or corrupted glyphs. |
| F-06 | Headings | **Do:** use plain-text canonical headings, larger/bold rather than graphical banners. **Do not:** place heading text inside shapes or use creative metaphors. | **Do:** use conventional scholarly headings and subheadings to separate status and contribution types. **Do not:** hide academic categories inside narrative prose. |
| F-07 | Headers and footers | **Do:** keep name/contact in the main body and avoid essential data in headers/footers. **Do not:** put the email, phone, or section content only in repeated page furniture. | **Do:** after page one, a modest `Name — Page N` header/footer can help humans collate pages; keep substantive data in the body. **Do not:** let page furniture collide with citations or entries. |
| F-08 | Tables and text boxes | **Do:** avoid tables and text boxes entirely in the ATS version. **Do not:** use invisible tables for alignment. | **Do:** use paragraph/tab structures for the main CV; a simple funder-mandated table is acceptable. **Do not:** impose the ATS ban over a required academic template, or use complex nested tables voluntarily. |
| F-09 | Images and icons | **Do:** omit photo, logos, QR codes, skill bars, charts, and icon-only contact labels. **Do not:** encode information visually without text. | **Do:** keep the CV text-led; include figures only when a specific portfolio/application format requests them. **Do not:** add decorative institutional logos, portraits, or charts to a standard CV. |
| F-10 | Bullets | **Do:** use standard round bullets and concise 1–2 line statements. **Do not:** use arrows, emoji, checkmarks, or custom glyphs. | **Do:** use bullets selectively for research responsibilities, findings, teaching scope, or service; full citations may remain unbulleted. **Do not:** force every academic entry into a corporate one-line bullet. |
| F-11 | Emphasis | **Do:** use bold consistently for employer/role hierarchy and italics sparingly. **Do not:** bold every keyword or underline long passages. | **Do:** emphasize applicant name in author lists and use discipline-standard italics/bold for venue/title hierarchy where helpful. **Do not:** distort bibliographic conventions for visual drama. |
| F-12 | Dates | **Do:** right-align through simple tabs if stable, or keep inline; use one format throughout. **Do not:** use `03/04/24` or omit years. | **Do:** use reverse chronology inside sections and one consistent convention. **Do not:** mix seasons, numeric dates, and years without reason. |
| F-13 | Links | **Do:** keep links short, visible, descriptive, clickable, and verified; prefer one strong portfolio URL to many noisy links. **Do not:** use raw tracking URLs. | **Do:** use persistent identifiers such as DOI and ORCID when available. **Do not:** overload each citation with redundant URLs or link to material the committee cannot access. |
| F-14 | Color | **Do:** use black/dark text on white; one restrained accent is safe only if contrast and grayscale printing remain strong. **Do not:** depend on color for category meaning. | **Do:** use a conservative, high-contrast palette, normally monochrome. **Do not:** let branding compete with scholarship or reduce print readability. |
| F-15 | Page breaks | **Do:** keep each employer/role header with at least one following bullet and avoid stranded headings. **Do not:** leave a nearly empty second page. | **Do:** prevent citations, entries, and section headings from splitting awkwardly; repeat name/page number. **Do not:** squeeze the CV to avoid natural page breaks. |
| F-16 | Filename | **Do:** use `FirstName_LastName_Resume_Role.pdf` or the employer’s requested convention. **Do not:** submit `resume_final_v7_REAL.pdf`. | **Do:** use `FirstName_LastName_CV_Program_or_Funder.pdf` when useful and permitted. **Do not:** include informal version labels or an incorrect institution name. |
| F-17 | Extraction test | **Do:** run plain-text extraction and confirm the order is name → contact → sections → entries; check for missing/duplicated text. **Do not:** validate only by visual inspection. | **Do:** run both visual and text/search checks, then inspect the portal preview. **Do not:** assume a visually correct PDF has correct searchable content. |
| F-18 | Accessibility | **Do:** maintain logical reading order, real text, sufficient contrast, and descriptive link text. **Do not:** create an ATS-safe file that is inaccessible to people. | **Do:** preserve heading hierarchy, readable contrast, Unicode text, and screen-readable ordering where possible. **Do not:** treat academic convention as an excuse for inaccessible layout. |

## 5. Matching strategy: corporate retrieval versus academic alignment

Corporate retrieval can combine literal Boolean search, normalized skill taxonomies, semantic matching, and employer-defined criteria. Textkernel explicitly distinguishes human-authored search from automated matching and allows category weighting; Workday’s Candidate Skills Match gives greater weight to required skills but, for that specific feature, does not consider skill recency, duration, or total work experience. These vendor-specific facts show why no universal density score should be reverse-engineered. Academic review instead evaluates the coherence and credibility of research preparation and fit. [Textkernel Search & Match FAQ](https://developer.textkernel.com/tx-platform/v10/faq/), [Workday Candidate Skills Match](https://doc.workday.com/admin-guide/en-us/human-capital-management/recruiting/candidates/candidate-skills-match/bmj1604095304483.html), [MIT personal-statement criteria](https://mitcommlab.mit.edu/broad/commkit/graduate-school-personal-statement/), [Cornell research-statement guidance](https://gradschool.cornell.edu/career-and-professional-development/pathways-to-success/prepare-for-your-career/take-action/research-statement/).

| ID | Matching insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| K-01 | No universal keyword-density target | **Do:** optimize truthful requirement coverage and evidence. **Do not:** target a percentage, repeat a term unnaturally, or promise an ATS score. | **Do:** optimize thematic coherence and evidence of research potential. **Do not:** calculate “research keyword density.” |
| K-02 | Requirement classification | **Do:** classify posting terms as: eligibility/knockout, required skill, preferred skill, responsibility, domain, outcome, and interpersonal criterion. **Do not:** treat every word in the posting as equally important. | **Do:** classify target signals as: eligibility, research question/theory, method, domain/data, output, PI/lab resources, teaching, and funding criteria. **Do not:** treat a program page as a bag of terms. |
| K-03 | Required versus preferred | **Do:** prioritize verified must-haves and required skills; then add preferred skills with evidence. **Do not:** crowd out core requirements with attractive but optional technologies. | **Do:** prioritize prerequisites and the central research fit; then show complementary methods or topics. **Do not:** overstate peripheral overlap while missing the lab’s central problem. |
| K-04 | Exact term plus evidence | **Do:** use the job’s exact standard term once in Technical Skills and/or an evidence-bearing bullet when true. **Do not:** add a synonym list disconnected from experience. | **Do:** use the field’s accepted term and demonstrate how it was applied in research. **Do not:** repeat fashionable terms without research substance. |
| K-05 | Synonyms and abbreviations | **Do:** pair common forms naturally—`Amazon Web Services (AWS)`, `CI/CD`, `PostgreSQL`—when they improve retrieval. **Do not:** insert every possible spelling variation. | **Do:** define non-universal acronyms at first use and retain the discipline’s preferred label. **Do not:** rename methods to mirror a lab if the method is only adjacent. |
| K-06 | Skill context | **Do:** support high-priority skills in Experience or Projects, not only a skills list. **Do not:** infer proficiency from a dependency file or course title alone. | **Do:** support methods/tools through a research entry, output, thesis, or teaching record. **Do not:** imply methodological independence from mere exposure. |
| K-07 | Role/title matching | **Do:** preserve the official title; a factual functional clarification in parentheses is allowed, e.g., `Software Engineer (Backend)`. **Do not:** rename a role to the target title. | **Do:** preserve official appointment titles and explain function separately. **Do not:** relabel `Research Assistant` as `Researcher`, `Instructor`, or `Principal Investigator`. |
| K-08 | Production impact | **Do:** prioritize reliability, latency, throughput, cost, revenue, quality, user/tenant scope, incident reduction, delivery speed, or risk when verified. **Do not:** invent business value from a technical artifact. | **Do:** prioritize research significance, methodological contribution, evaluation quality, findings, reproducibility, scholarly output, and future potential. **Do not:** force every research outcome into revenue or efficiency language. |
| K-09 | Modern system design | **Do:** show constraints, components, interfaces, data flow, tradeoffs, failure modes, security, observability, deployment, and operational ownership where evidenced. **Do not:** call routine CRUD work “distributed systems architecture.” | **Do:** show assumptions, theoretical basis, research design, baselines, evaluation protocol, findings, limitations, and reproducibility. **Do not:** market an architecture as a research contribution without novelty/evaluation evidence. |
| K-10 | Agile/delivery methods | **Do:** mention Scrum, Kanban, CI/CD, code review, testing, incident response, or release practice only when actually used and relevant. **Do not:** list ceremonies as impact or infer Scrum from Jira usage. | **Do:** mention project management practices only when material to collaborative research delivery. **Do not:** let agile vocabulary replace research methodology. |
| K-11 | Research alignment | **Do:** for research-engineering roles, connect domain problem, method, implementation, and deployment/evaluation. **Do not:** present a publication list without job relevance. | **Do:** establish alignment through question + method + domain + prior result/output + plausible future direction + specific lab/program capability. **Do not:** use faculty-name repetition as a substitute. |
| K-12 | Theoretical foundations | **Do:** include theory only when it explains engineering decisions or is requested. **Do not:** make a production resume read like a literature review. | **Do:** name relevant theoretical frameworks precisely and show how they guided a method or interpretation. **Do not:** claim theoretical grounding from a citation list alone. |
| K-13 | Output assessment | **Do:** select outputs that prove the target capability: deployed systems, maintained libraries, patents, standards, or relevant papers. **Do not:** optimize for raw output count. | **Do:** preserve venue, authorship, status, contribution, and significance; follow discipline norms where conferences, journals, software, or datasets carry different weight. **Do not:** treat all outputs as equivalent. |
| K-14 | Recency and trajectory | **Do:** favor recent, relevant evidence while retaining older exceptional work. **Do not:** hide a current role beneath older projects. | **Do:** show development from preparation to independent questions and future direction. **Do not:** reorder chronology so aggressively that the scholarly trajectory becomes misleading. |
| K-15 | Soft-skill evidence | **Do:** demonstrate collaboration, leadership, mentoring, stakeholder work, and communication through actions and scope. **Do not:** list `team player`, `excellent communicator`, or `leadership` without evidence. | **Do:** evidence mentoring, interdisciplinary collaboration, teaching, peer review, service, and communication through named activities. **Do not:** use personality claims as proof of academic citizenship. |

### 5.1 Generator-side matching matrices

These are **content-selection tools**, not claims about proprietary ATS formulas.

**Corporate evidence matrix**

| Target class | Evidence question | Include when |
|---|---|---|
| Eligibility | Does the verified profile meet a non-negotiable requirement? | Answer only in the appropriate application field or resume when explicitly useful. Never infer. |
| Required skill | Is there direct, dated evidence of use? | Direct evidence exists; place the exact standard term naturally. |
| Preferred skill | Is there direct or clearly adjacent evidence? | Direct evidence exists, or describe the adjacent capability without claiming the requested skill. |
| Responsibility | Has the candidate performed comparable work? | A role/project proves the action, scope, and ownership. |
| Domain | Has the candidate worked in the target problem space? | The domain is verified; otherwise describe transferable technical evidence, not false domain experience. |
| Outcome | Is a result measurable or concretely observable? | Source supports the metric or a precise qualitative outcome. |

**Academic alignment matrix**

| Target class | Evidence question | Include when |
|---|---|---|
| Research question/theory | Has the applicant worked on the same or a logically connected problem? | The connection can be explained, not merely keyword-matched. |
| Method | Has the method been applied, evaluated, taught, or published? | The level of experience can be stated precisely. |
| Domain/data/system | Is there direct knowledge of the target setting? | Evidence exists; otherwise state a transferable setting without claiming domain expertise. |
| Output | What verifiable artifact resulted? | Status, authorship, venue, and link/identifier are correct. |
| Independence | What was the applicant’s own intellectual/technical contribution? | Sources or the user can distinguish it from team/PI work. |
| Future fit | Can prior evidence support a plausible next step in the target lab/program? | The connection names a real capability/resource and avoids promising a fixed thesis prematurely. |

## 6. Canonical taxonomy and section labels

Standard headings improve both parser mapping and human navigation. They are safe output defaults—not magic tokens. Empty sections must be omitted, and explicit portal/discipline requirements override the order.

### 6.1 Corporate canonical order

1. Name and Contact Information (unlabeled)
2. Professional Summary *(optional; only if tailored and evidence-based)*
3. Technical Skills
4. Professional Experience
5. Selected Projects *(when they add evidence not already shown)*
6. Education
7. Certifications *(if relevant)*
8. Selected Publications / Open Source / Awards / Leadership *(only when role-relevant)*

### 6.2 Academic canonical order

1. Name and Contact Information (unlabeled)
2. Research Interests *(specific, optional for established scholars or when redundant)*
3. Education
4. Research Experience
5. Publications, separated by verified status/type
6. Conference Presentations and Invited Talks
7. Teaching Experience
8. Grants and Fellowships
9. Honors and Awards
10. Industry Experience *(if relevant)*
11. Professional Service
12. Technical Skills / Research Methods / Languages
13. Professional Memberships *(if meaningful)*
14. References *(only when requested/customary and consented)*

### 6.3 Ivy League, Oxbridge, and research-funder modes

“Academic CV” is not one fixed form. Graduate-admissions CVs, faculty-job CVs, PI-outreach documents, narrative CVs, and funder biosketches have different constraints. Princeton describes a graduate-student CV as a living record with no inherent length limit, typically two or three pages at that stage, and recommends reordering headings for the reader. Oxford likewise distinguishes the no-fixed-page-limit academic CV from the shorter non-academic CV and asks applicants to tailor it to the research/teaching role. These are audience principles, not permission to ignore a portal template. [Princeton graduate CV guide](https://careerdevelopment.princeton.edu/guides/resume-cv-cover-letter-perspective-statement/cv-writing-guide-for-graduate-students), [Yale academic CV guide](https://ocs.yale.edu/%F0%9F%93%84-your-academic-cv/), [Oxford academic CV guidance](https://www.careers.ox.ac.uk/cvs), [Cambridge PhD/postdoc CV guide](https://www.careers.cam.ac.uk/files/phdpostdoccvbook.pdf).

| ID | Institution / scheme insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| AC-01 | Ivy/Oxbridge guidance is contextual, not a universal template | **Do:** use university career guidance for readability and evidence principles when relevant. **Do not:** turn a graduate academic CV example into a corporate software-resume template. | **Do:** tailor section order and depth to admissions, research, or teaching emphasis. **Do not:** label one Princeton, Yale, Oxford, or Cambridge example as the single “Ivy/Oxbridge standard.” |
| AC-02 | Comprehensive does not mean indiscriminate | **Do:** select only role-relevant evidence within the requested 1–2 pages. **Do not:** cite “CVs have no length limit” to justify a long corporate resume. | **Do:** preserve the complete relevant scholarly record and prune repetition, stale detail, and irrelevant employment. **Do not:** confuse exhaustiveness with verbosity; an early-stage graduate CV may naturally be only two or three pages. |
| AC-03 | Research/teaching/service ordering is a signal | **Do:** order experience by relevance while keeping chronology inside sections. **Do not:** lead with teaching or publications unless they support the role. | **Do:** put Research Experience/Publications earlier for research-heavy selection and Teaching Experience earlier for teaching-heavy selection; preserve accurate chronology within each section. **Do not:** use one fixed order for all academic applications. |
| AC-04 | PI/lab outreach requires specific intellectual fit | **Do:** for an industrial research role, connect the employer's problem to evidenced methods and outputs. **Do not:** praise the organization generically. | **Do:** connect a real PI paper/problem/resource to the applicant's prior question, method, or credible next step. **Do not:** mass-produce faculty name-drops or claim fit from topic overlap alone. [Oxford supervisor guidance](https://www.careers.ox.ac.uk/approaching-phd-supervisors). |
| AC-05 | ERC 2026 is a bounded CV-and-track-record mode | **Do:** if translating an ERC record for industry, select the few outputs that prove the target capability. **Do not:** copy a grant-form bibliography into a software resume. | **Do:** for a 2026 ERC call, use the required merged CV and Track Record structure, observe the four-page limit, select up to ten outputs that best evidence contribution to knowledge, and explain the applicant's contribution/independence factually. **Do not:** submit a general unlimited CV or rank outputs solely by count or venue prestige. [ERC 2026 call guidance](https://erc.europa.eu/news-events/events/erc-grants-what-expect-2026-calls). |
| AC-06 | UKRI R4RI is a narrative evidence mode | **Do:** borrow its contribution-evidence discipline only when useful; keep the corporate resume concise. **Do not:** paste long narrative modules into an industry resume. | **Do:** use the current call's R4RI modules to evidence a broad range of contributions, outcomes, roles, and context. **Do not:** substitute a publication-list CV or assume every UKRI opportunity uses identical limits. [UKRI R4RI guidance](https://www.ukri.org/apply-for-funding/develop-your-application/resume-for-research-and-innovation-r4ri-guidance/). |
| AC-07 | DFG requires its own template | **Do:** keep DFG-form content separate from a corporate resume and translate only verified, relevant evidence. **Do not:** assume a DFG-formatted dossier is recruiter-ready. | **Do:** use the mandatory current DFG CV template and the exact programme/call instructions; where requested, limit the selected publication list as specified. **Do not:** reformat the DFG submission into a generic academic CV. [DFG CV FAQ](https://www.dfg.de/en/research-funding/proposal-funding-process/faq/cv), [DFG proposal information](https://www.dfg.de/en/research-funding/proposal-funding-process/individual-grants-programmes/proposal-information). |
| AC-08 | NSF and NIH use certified structured forms | **Do:** treat government biosketches as source records, not as a corporate resume format. **Do not:** convert disclosures or support records into marketing claims. | **Do:** use SciENcv for NSF-required documents and, for NIH submissions on or after 8 May 2026, the compliant Common Forms plus NIH supplement as applicable; preserve certification, ORCID, version, and disclosure accuracy. **Do not:** upload a hand-built generic CV where a digitally certified form is required. [NSF senior-personnel documents](https://www.nsf.gov/funding/senior-personnel-documents), [NSF SciENcv FAQ](https://www.nsf.gov/policies/document/faq-using-sciencv), [NIH enforcement notice](https://grants.nih.gov/grants/guide/notice-files/NOT-OD-26-079.html). |
| AC-09 | Live call instructions are the authority | **Do:** snapshot the current job description and submission requirements before generation. **Do not:** let any university or vendor guide override the employer's explicit instructions. | **Do:** snapshot the current call, programme, portal, and template version before generation. **Do not:** rely on a prior-year page limit, an example CV, or remembered funder rule. |

| ID | Labeling insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| L-01 | Opening profile | **Do:** use `Professional Summary` only for a 2–3 line role-specific value proposition grounded in evidence. **Do not:** use `Objective`, generic adjectives, or a paragraph biography. | **Do:** use `Research Interests` for a concise, specific set of coherent questions/domains/methods when useful. **Do not:** use a corporate summary or list every fashionable topic. |
| L-02 | Work history | **Do:** use `Professional Experience`. **Do not:** use `Career Journey`, `Employment Story`, or `What I’ve Done`. | **Do:** use `Research Experience`, `Academic Appointments`, and/or `Industry Experience` as factually appropriate. **Do not:** merge all roles under a vague `Experience` heading if scholarly status becomes unclear. |
| L-03 | Projects | **Do:** use `Selected Projects` and include only role-relevant work. **Do not:** use `Portfolio` as a catch-all for unverified links. | **Do:** place research projects in `Research Experience`; use `Selected Research Projects` only when needed for early-stage applicants. **Do not:** make completed research look like a hobby project. |
| L-04 | Skills | **Do:** use `Technical Skills` with typed categories. **Do not:** use `Toolbox`, `Superpowers`, `Expertise Cloud`, or rating graphics. | **Do:** use `Research Methods`, `Technical Skills`, and/or `Languages` depending on the field. **Do not:** collapse methods, programming languages, and spoken languages into one undifferentiated list. |
| L-05 | Education | **Do:** use `Education`. **Do not:** use `Academic Background` if a parser-friendly standard label is available. | **Do:** use `Education`; add thesis/dissertation details within entries. **Do not:** hide degree status in prose. |
| L-06 | Publications | **Do:** use `Selected Publications` only when relevant; otherwise omit. **Do not:** label blogs or internal reports as peer-reviewed publications. | **Do:** use `Publications` plus explicit subheadings such as `Peer-Reviewed Journal Articles`, `Peer-Reviewed Conference Papers`, `Preprints`, and `Manuscripts Under Review`, following field norms. **Do not:** put `in preparation` work among published items. |
| L-07 | Talks and presentations | **Do:** use `Selected Talks` only for role-relevant speaking evidence. **Do not:** confuse attendance with presentation. | **Do:** use `Invited Talks`, `Conference Presentations`, and `Posters` accurately. **Do not:** classify a poster, panel attendance, or non-archival workshop as a journal/conference publication. |
| L-08 | Teaching | **Do:** place relevant training/mentoring under the role or `Leadership and Mentoring`; use `Teaching Experience` only when central. **Do not:** over-expand routine onboarding. | **Do:** use `Teaching Experience` and state course, institution, role, level, dates, responsibilities, and independently taught scope. **Do not:** call grading/support duties full course instruction. |
| L-09 | Funding and recognition | **Do:** use `Awards` or `Certifications` only for verified, relevant items. **Do not:** combine pending applications with received awards. | **Do:** separate `Grants and Fellowships` from `Honors and Awards`; identify role and status. **Do not:** imply the applicant was PI on a grant merely because they worked on the funded project. |
| L-10 | Service | **Do:** use `Leadership and Community` only when it supports the target role. **Do not:** include every club membership. | **Do:** use `Professional Service` for reviewing, committees, conference organization, societies, mentoring/service, as discipline-appropriate. **Do not:** inflate membership into leadership. |
| L-11 | References | **Do:** omit the section unless requested. **Do not:** write `References available upon request`. | **Do:** use `References` only when requested or conventional, with name, title, institution, relationship/context if useful, email, and consent. **Do not:** list private phone numbers without permission. |
| L-12 | Empty/weak categories | **Do:** omit an empty or low-value section. **Do not:** create `Publications`, `Patents`, or `Awards` headings with no entries. | **Do:** omit empty sections and elevate strong research/teaching evidence. **Do not:** write `None` or invent works in progress to fill perceived gaps. |

## 7. Zero-embarrassment truth and hallucination guardrails

### 7.1 Claim-state model

Every candidate fact must carry an internal state before generation:

| State | Definition | Allowed in a submission-ready document? |
|---|---|---|
| **Verified** | Supported by a primary artifact, official record, publication/DOI, repository evidence, certificate, analytics, or an unambiguous user-confirmed fact. | Yes, with faithful wording. |
| **User-asserted** | Explicitly supplied by the candidate but not independently verified. | Yes when plausible and non-conflicting; high-risk claims should be confirmed before finalization. |
| **Derived** | Calculated from complete verified inputs using a transparent rule. | Yes only when the calculation is valid, useful, and its meaning is not overstated. |
| **Inferred** | Suggested by context, code, a dependency, title, or neighboring facts but not directly established. | No. Use only to ask a clarification question. |
| **Unknown** | Missing or ambiguous. | No. Omit or ask; never fill from model knowledge. |
| **Conflicted** | Sources disagree. | No. Surface the conflict and resolve it. |

| ID | Truth-control insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| G-01 | Source-bounded generation | **Do:** rewrite only facts present in the evidence ledger. **Do not:** use the model’s knowledge of a company, stack, or typical role to fill gaps. | **Do:** rewrite only documented scholarly facts. **Do not:** infer a method, finding, venue status, or contribution from a paper title or field norm. |
| G-02 | Unknown is not zero or no | **Do:** treat an absent metric, team size, technology, or outcome as unknown. **Do not:** output `0`, `none`, or a negative claim unless verified. | **Do:** treat absent publications, funding, teaching, or awards as unknown. **Do not:** state `No publications` or `No funding` unless the user explicitly wants and confirms it. |
| G-03 | Metrics | **Do:** use an exact number only when the source defines the measure, baseline, endpoint, unit, and relevant time window. **Do not:** manufacture round percentages or “estimated” impact. | **Do:** use sample size, dataset size, performance measure, award amount, student count, or citation data only from valid sources and with context. **Do not:** convert a qualitative finding into a numeric claim. |
| G-04 | Percentage claims | **Do:** verify numerator, denominator, baseline, comparison, and period; preserve `approximately`, `up to`, median, or average exactly. **Do not:** turn `up to 30%` into `30%` or combine unrelated improvements. | **Do:** preserve confidence intervals, uncertainty, significance, and evaluation conditions when relevant. **Do not:** turn correlation into causation or a limited result into a general effect. |
| G-05 | Technology claims | **Do:** require direct evidence of use or explicit user confirmation before listing a language/framework/cloud service. **Do not:** infer use from a neighboring project, organization stack, generated lockfile, or one mention in documentation. | **Do:** distinguish implementation use, experimental use, coursework, and conceptual familiarity. **Do not:** call exposure expertise. |
| G-06 | Proficiency | **Do:** prefer evidence over `expert`, `advanced`, or years-based ratings. **Do not:** assign proficiency levels automatically. | **Do:** describe applied methods and outputs. **Do not:** label the applicant an expert based on one project or paper. |
| G-07 | Ownership verbs | **Do:** use `led`, `owned`, `architected`, or `mentored` only when decision authority, responsibility, or coordination is explicit. **Do not:** infer leadership from seniority, commit count, or being the only named person. | **Do:** state exact intellectual/experimental contribution and role. **Do not:** infer independence, first-author leadership, supervision, or PI responsibility. |
| G-08 | Production claims | **Do:** reserve `production`, `deployed`, `operated`, `on-call`, `high availability`, and `SLA` for supported operational evidence. **Do not:** call a demo, prototype, testnet, or architecture proposal production. | **Do:** label prototype, proof of concept, simulation, pilot, deployed system, and field study precisely. **Do not:** collapse them into an implemented real-world system. |
| G-09 | Quality/security/scalability | **Do:** support `secure`, `scalable`, `fault-tolerant`, `compliant`, or `optimized` with architecture, controls, tests, benchmarks, audits, or operational data. **Do not:** use these as default adjectives. | **Do:** support validity, robustness, reproducibility, privacy, ethics approval, or generalizability with the relevant design/evidence. **Do not:** claim them from intent alone. |
| G-10 | Research findings | **Do:** for industry R&D, state a finding only when a completed evaluation supports it. **Do not:** present a roadmap or hypothesis as achieved impact. | **Do:** distinguish research question, hypothesis, method, preliminary observation, result, interpretation, and limitation. **Do not:** use `proved` when the work only suggests, supports, or demonstrates under stated conditions. |
| G-11 | Publication status | **Do:** label a role-relevant item as published, accepted, forthcoming, preprint, under review, or in preparation exactly. **Do not:** use `publication` for an unsubmitted manuscript. | **Do:** place accepted/forthcoming work with publications only when acceptance is real; separate preprints, under-review manuscripts, and works in progress. **Do not:** upgrade status. Berkeley explicitly warns that `forthcoming` requires actual acceptance. |
| G-12 | Authorship | **Do:** preserve author order and contribution statement if included. **Do not:** imply sole authorship because only the candidate is discussed in the bullet. | **Do:** reproduce the complete citation and exact author order; bold the applicant’s name if the field permits. **Do not:** reorder authors or infer equal/first/corresponding authorship without evidence. |
| G-13 | Funding and grants | **Do:** state funded-project participation separately from grant ownership. **Do not:** claim `secured $X` unless the candidate’s role and awarded amount are verified. | **Do:** state funder, scheme, project, role (PI/co-PI/team member/recipient), status, period, and verified amount as appropriate. **Do not:** call submitted/pending funding awarded. |
| G-14 | Teaching roles | **Do:** translate teaching to mentoring/training only when relevant and accurate. **Do not:** inflate course assistance into organizational leadership. | **Do:** distinguish Instructor of Record, Lecturer, Teaching Assistant, lab demonstrator, grader, and guest lecturer. **Do not:** merge their authority or responsibility. |
| G-15 | Dates and current status | **Do:** use `Present` only after explicit current-status confirmation or unambiguous source evidence. **Do not:** assume the latest role continues. | **Do:** distinguish completed, current, expected, scheduled, and forthcoming dates. **Do not:** present an expected degree or future talk as completed. |
| G-16 | Confidentiality | **Do:** abstract confidential client/system details while retaining truthful scope, e.g., `European financial-services client`, when authorized. **Do not:** reveal protected names, architecture, metrics, vulnerabilities, or customer data. | **Do:** respect embargoes, anonymization, participant confidentiality, unpublished results, and collaboration agreements. **Do not:** expose restricted datasets, reviewer identities, or confidential findings. |
| G-17 | Conflicting sources | **Do:** stop and ask when titles, dates, metrics, or roles conflict; record the resolved source. **Do not:** silently choose the more impressive version. | **Do:** stop on conflicts in author order, status, venue, degree, supervisor, grant role, or dates. **Do not:** reconcile scholarly records through guesswork. |
| G-18 | Cross-document consistency | **Do:** align the final resume with application fields, LinkedIn, portfolio, and interview facts, allowing only truthful differences in emphasis. **Do not:** create contradictory titles/dates or a skill that exists only in one generated version. | **Do:** align CV, statement, portal, transcript, publication record, ORCID, and referee information. **Do not:** let tailoring change facts or the claimed research trajectory. |
| G-19 | Draft versus final mode | **Do:** in draft mode, use explicit `NEEDS USER VERIFICATION` notes outside candidate-facing text; in final mode, omit unresolved claims. **Do not:** hide uncertainty. | **Do:** maintain the same mode separation, especially for publication, funding, authorship, and finding status. **Do not:** export internal confidence labels into the submitted CV. |
| G-20 | No deceptive ATS tactics | **Do:** reject white text, hidden layers, tiny keyword blocks, copied job descriptions, metadata stuffing, or irrelevant keyword repetition. **Do not:** help the candidate misrepresent fit. | **Do:** keep the PDF honest and readable. **Do not:** insert hidden faculty names, paper titles, or field terms to manipulate search. |

Publication-status guidance is supported by [Berkeley’s academic CV guidance](https://career.berkeley.edu/grad-students-postdocs/academic-job-search/the-cv-part-2-elements/), which separates accepted/forthcoming work from submitted or under-review work.

## 8. Tone, vocabulary, and anti-AI-signature controls

Harvard advises résumé language that is specific, active, clear, fact-based, and unembellished; it also notes that generative AI should not become the primary author because the output tends to be generic. Academic guidance from Cornell similarly favors clear, concrete, readable writing and warns against unnecessary jargon. [Harvard resume guidance](https://careerservices.fas.harvard.edu/resources/hes-create-impactful-resumes-and-cover-letters/), [Cornell research-statement guidance](https://gradschool.cornell.edu/career-and-professional-development/pathways-to-success/prepare-for-your-career/take-action/research-statement/).

| ID | Tone insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| T-01 | Voice | **Do:** use concise, active, results-aware fragments without first-person pronouns. **Do not:** write a promotional narrative or self-evaluation. | **Do:** use precise, objective, field-aware language; concise explanatory sentences are acceptable. **Do not:** imitate corporate hype or remove necessary methodological nuance. |
| T-02 | Verb selection | **Do:** prefer plain verbs that identify the work: built, designed, implemented, deployed, migrated, reduced, automated, tested, diagnosed, maintained, led, mentored. **Do not:** choose a grander verb than the evidence supports. | **Do:** prefer investigated, analyzed, developed, derived, evaluated, compared, characterized, validated (only when valid), taught, supervised, presented, reviewed. **Do not:** use verbs that exaggerate epistemic certainty. |
| T-03 | Claims versus adjectives | **Do:** replace `robust platform` with the concrete mechanism/result that made it reliable. **Do not:** use adjectives as evidence. | **Do:** replace `novel and groundbreaking method` with the specific difference, comparison, and result. **Do not:** claim novelty without a defensible prior-art basis. |
| T-04 | Jargon | **Do:** use industry terms needed by the role and explain internal acronyms. **Do not:** stack buzzwords to sound senior. | **Do:** use necessary disciplinary terminology while keeping the entry intelligible to adjacent-field committee members. **Do not:** oversimplify away the method or bury meaning in jargon. |
| T-05 | Energy level | **Do:** create energy through specificity, ownership, scale, and outcomes. **Do not:** use exclamation, motivational language, or self-praise. | **Do:** create authority through clarity, evidence, coherent trajectory, and careful status labels. **Do not:** use sales intensity as a proxy for research significance. |
| T-06 | Repetition | **Do:** vary verbs only when the facts differ; repeated accurate `built` is better than thesaurus inflation. **Do not:** force a unique dramatic verb for every bullet. | **Do:** repeat standard methodological terms where accuracy requires it. **Do not:** replace precise terms merely for stylistic variation. |

### 8.1 Hard-ban lexicon for generated prose

Ban these by default unless they are part of an official title, proper noun, direct quotation, or technically necessary phrase:

`delve`, `delved`, `testament`, `spearheaded`, `leveraged` (when it only means “used”), `utilized`, `harnessed`, `revolutionized`, `game-changing`, `groundbreaking`, `cutting-edge`, `world-class`, `best-in-class`, `visionary`, `dynamic professional`, `results-driven`, `detail-oriented`, `team player`, `go-getter`, `thought leader`, `proven track record`, `synergy`, `passionate about`, `seamlessly`, `successfully` (when the result already proves success), `significantly` (without a supported statistical or quantified meaning), `impactful`, `transformative`, `innovative` (without stating what is new), `robust` (without a defined property), `responsible for`, `helped with`, `worked on`, `various`, `etc.`

| ID | Banned-language application | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| T-07 | Weak duty phrases | **Do:** replace `responsible for APIs` with the verified action and scope. **Do not:** leave ownership vague. | **Do:** replace `worked on research` with the question, method, and contribution. **Do not:** imply contribution without stating it. |
| T-08 | Unsupported superlatives | **Do:** use a benchmark, scale, or comparison if verified. **Do not:** write `best`, `leading`, `world-class`, or `state-of-the-art` without a defensible source. | **Do:** use `state-of-the-art` only in a properly supported research context. **Do not:** self-certify novelty, excellence, or field leadership. |
| T-09 | Epistemic precision | **Do:** separate observed system outcome from inferred business effect. **Do not:** say a feature `drove revenue` unless attribution is supported. | **Do:** use `suggests`, `supports`, `is consistent with`, `demonstrates under X conditions`, or `establishes` according to evidence. **Do not:** convert association into proof. |
| T-10 | Human voice | **Do:** preserve the candidate’s natural terminology and explain domain-specific work plainly. **Do not:** homogenize every candidate into the same action-verb template. | **Do:** preserve disciplinary voice and accurate technical language. **Do not:** make every project description sound like an abstract generated from one template. |

### 8.2 2026 AI-assisted application and fraud detection: documented reality

The generator must not promise to “beat AI detection.” Public 2026 product documentation shows a more concrete landscape:

- Greenhouse's Fraud Detection uses device and contact signals such as IP address, user agent, email/phone traits, location/time-zone mismatch, and related enrichment. Greenhouse says the feature is not generative AI, produces a report for recruiter review, and does not itself make an automated rejection decision. Its organization-managed spam blocklist is a separate intake control. [Greenhouse security/privacy FAQ](https://support.greenhouse.io/hc/en-us/articles/45397259312027-Fraud-Detection-and-Spam-Blocklist-Security-Privacy-FAQ), [Greenhouse fraud policy guide](https://support.greenhouse.io/hc/en-us/articles/44681941657243-Operational-readiness-guide-Fraud-Detection-policy).
- Greenhouse and Lever document identity verification, fraud signals, résumé/profile inconsistencies, deepfake or substituted interviewees, and inability to substantiate claimed skills as integrity concerns. Lever also explicitly distinguishes assistance with structure or wording from AI replacing a candidate's experience or ability. [Greenhouse identity verification](https://support.greenhouse.io/hc/en-us/articles/40966215931291-Greenhouse-identity-verification-overview), [Lever 2026 candidate-fraud guidance](https://www.lever.co/blog/what-the-rise-in-ai-powered-candidate-fraud-really-means-for-ta-teams).
- NIST continues to evaluate text discriminators and reports highly system-dependent results: some generators deceive most discriminators while some discriminators identify output from almost all tested generators. That is benchmark evidence, not proof that mainstream ATS products universally classify résumé prose as human- or AI-written. [NIST text-to-text evaluation](https://www.nist.gov/publications/2024-nist-genai-pilot-study-text-text-evaluation-overview-and-results).

| ID | Integrity insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| AI-01 | No documented universal “AI résumé detector” | **Do:** assume employers may use matching, fraud, identity, assessment, and human-review tools. **Do not:** claim that Workday, Greenhouse, Lever, Taleo, or iCIMS universally rejects AI-written text, or optimize against an imaginary detector score. | **Do:** comply with each institution's disclosed AI-use policy and preserve personal authorship of ideas and claims. **Do not:** claim every admissions portal runs a reliable AI-writing detector. |
| AI-02 | Fraud signals are not prose-style scores | **Do:** keep identity, contact, location, dates, and profile facts consistent and truthful. **Do not:** provide advice for manipulating IP, device, email, phone, location, spam, liveness, or identity-verification signals. | **Do:** keep CV, portal, transcript, publication record, and identity data consistent. **Do not:** treat authentication or research-integrity checks as wording problems to evade. |
| AI-03 | AI assistance versus misrepresentation | **Do:** use AI to organize and edit source-bounded facts while ensuring the candidate can substantiate every claim. **Do not:** let AI invent experience, answer assessments for the candidate, impersonate them, or replace demonstrated skill. | **Do:** use AI only within the programme's policy and retain the applicant's intellectual ownership and ability to defend the work. **Do not:** generate findings, citations, research ideas presented as personal work, or undisclosed prohibited material. |
| AI-04 | Human suspicion patterns are not deterministic detectors | **Do:** remove generic buzzwords, placeholder text, near-verbatim job-description copying, unsupported large numbers, and breadth without technical depth. **Do not:** label any one phrase as proof of AI authorship or intentionally add errors to look human. | **Do:** remove generic praise, fabricated specificity, citation errors, abrupt voice shifts, and claims the applicant cannot explain. **Do not:** equate polished or non-native English with misconduct. |
| AI-05 | Cross-stage consistency is the strongest defense | **Do:** ensure résumé, form, public profile, portfolio, references, screening answers, and interview explanations agree on the underlying facts. **Do not:** create tailored variants with contradictory dates, titles, stack, scale, or ownership. | **Do:** ensure the CV, statement, research proposal, transcript, ORCID, publications, referee account, and interview discussion agree. **Do not:** let tailoring alter status, authorship, contribution, or research trajectory. |
| AI-06 | Claim depth must survive follow-up | **Do:** retain an internal evidence packet for each major bullet: what, why, role, design choice, tradeoff, validation, and outcome. **Do not:** use impressive claims the candidate cannot explain under unscripted technical questioning. | **Do:** retain method, data, contribution, result/status, limitation, and output evidence for each research claim. **Do not:** present a method or finding the applicant cannot defend to a PI or panel. |
| AI-07 | Stylometric gaming is prohibited | **Do:** improve specificity, sentence variety, and candidate voice for readability. **Do not:** add typos, awkwardness, invisible characters, homoglyphs, paraphrase noise, or adversarial formatting to evade detection. | **Do:** preserve clear disciplinary prose and accessibility. **Do not:** degrade language or manipulate document encoding to influence an AI detector. |
| AI-08 | A risk flag is not proof | **Do:** design for truthful review and human appeal where available. **Do not:** describe a fraud signal, match category, or detector output as a conclusive hiring decision unless the employer's actual policy establishes that consequence. | **Do:** distinguish portal validation, misconduct review, committee judgment, and final decision. **Do not:** infer rejection reasons from an opaque flag or score. |
| AI-09 | Auditability beats “humanization” | **Do:** log source, claim state, transformation, and final wording for every substantive statement. **Do not:** run unsupported “humanizer” passes that break evidence traceability. | **Do:** retain source/citation/status provenance and any required AI-use disclosure. **Do not:** prioritize detector evasion over attribution, research integrity, or institutional policy. |

## 9. Portfolio, personal-site, and GitHub extraction rules

The model must extract a **fact ledger first** and generate prose second. A website is a claim source; a repository is an artifact source; neither alone establishes every business or research claim.

### 9.1 Evidence-source priority

1. Official records: degree/certificate/grant records, publisher/DOI pages, employer-authorized metrics, release/monitoring reports.
2. Explicit user-confirmed facts and original role/project records.
3. Repository artifacts: source files, manifests, tests, CI/CD, infrastructure, releases, issues/PRs, contribution history, license.
4. Personal website/portfolio and README claims.
5. Third-party profiles or summaries.
6. Model inference — never eligible for candidate-facing claims.

### 9.2 Canonical project fact schema

| Field | Required interpretation |
|---|---|
| `project_name` | Exact public or internal-safe name; record aliases for deduplication. |
| `context` | Employer/client/course/lab/open-source/personal; confidential handling. |
| `dates` | Verified start/end/current state; source. |
| `candidate_role` | Official role and exact individual contribution. |
| `problem_or_question` | Business/engineering problem or research question; do not invent. |
| `artifact` | Service, library, model, protocol, dataset, paper, experiment, course material, etc. |
| `technologies_or_methods` | Directly evidenced tools, methods, theories, datasets, and level of use. |
| `architecture_or_design` | Implemented decisions versus proposals; tradeoffs and constraints. |
| `scale` | Users, requests, data, chains, services, team, students, samples, experiments—only verified. |
| `validation` | Tests, benchmarks, monitoring, audit, peer review, experiment, ablation, user study, deployment. |
| `outcome_or_finding` | Exact supported result, qualitative outcome, or current status. |
| `output_status` | Deployed/prototype; published/accepted/preprint/under review/in preparation; released/private. |
| `evidence_pointer` | URL/file/line/record/user confirmation and confidence state. |

| ID | Extraction insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| E-01 | Separate extraction from writing | **Do:** build and validate the fact ledger before selecting bullets. **Do not:** summarize raw pages directly into impressive prose. | **Do:** build the same ledger with explicit method, contribution, status, and output fields. **Do not:** draft research claims from a homepage blurb. |
| E-02 | Website marketing claims | **Do:** treat `enterprise-grade`, user counts, performance, clients, and outcomes as user-asserted until verified. **Do not:** upgrade marketing copy to fact. | **Do:** treat `novel`, `peer-reviewed`, `funded`, `first`, and findings as status-sensitive claims requiring primary evidence. **Do not:** rely on lab/profile marketing text alone. |
| E-03 | Repository technology detection | **Do:** require meaningful source/config/test evidence of use; a manifest entry alone proves dependency presence, not proficiency or ownership. **Do not:** list every transitive package. | **Do:** distinguish code used to conduct research from methods the applicant designed or understood independently. **Do not:** treat a cloned research stack as methodological contribution. |
| E-04 | Contribution attribution | **Do:** use commits, PRs, issues, CODEOWNERS, release notes, and user confirmation to establish contribution; use `contributed to` when ownership is partial. **Do not:** infer lead ownership from repository visibility. | **Do:** combine contribution evidence with author/contributor statements and user confirmation. **Do not:** infer intellectual leadership from commit count or author position alone. |
| E-05 | Public versus open source | **Do:** call a repository open source only if it has a compatible license; otherwise say `public repository`. **Do not:** equate public visibility with open-source status. | **Do:** label research software/data license and availability accurately. **Do not:** call restricted or unlicensed artifacts open. |
| E-06 | Prototype versus production | **Do:** look for deployment configuration, releases, monitoring, operational docs, traffic/tenant evidence, and explicit confirmation. **Do not:** infer production from Docker, Terraform, or a live demo alone. | **Do:** distinguish conceptual design, prototype, controlled PoC, pilot, field deployment, and sustained operation. **Do not:** collapse these stages. |
| E-07 | Metrics extraction | **Do:** trace every number to analytics, reports, benchmarks, tickets, or direct user confirmation; retain definition and date. **Do not:** derive impact from stars, lines of code, commits, or repository age. | **Do:** trace sample size, accuracy, statistical values, citations, students, grant amounts, and experiment counts to primary evidence. **Do not:** use repository popularity as research impact. |
| E-08 | Architecture extraction | **Do:** distinguish implemented components from README plans, open issues, or future roadmap. **Do not:** claim the intended architecture as delivered. | **Do:** distinguish proposed method/protocol from implemented/evaluated contribution. **Do not:** report future work as current results. |
| E-09 | Date extraction | **Do:** use explicit project/employment records; repository timestamps are supporting evidence only. **Do not:** equate first/last commit with official project duration. | **Do:** use institutional, submission, conference, and publication records for dates/status. **Do not:** infer research duration from file history. |
| E-10 | Duplicate reconciliation | **Do:** merge website, GitHub, and role descriptions into one project identity while preserving source-specific evidence. **Do not:** count the same project as multiple achievements. | **Do:** connect a thesis, preprint, conference paper, dataset, and codebase without falsely treating them as independent research projects. **Do not:** inflate output count through duplicate manifestations. |
| E-11 | Broken/confidential links | **Do:** verify public accessibility; omit or replace private links with an authorized description. **Do not:** expose tokens, internal URLs, or client-confidential assets. | **Do:** use persistent public records or clearly label materials available on request when permitted. **Do not:** link embargoed manuscripts or restricted participant data. |
| E-12 | Missing facts | **Do:** ask targeted questions for role, scale, outcome, dates, or ownership when material. **Do not:** fill the gap with typical software-project assumptions. | **Do:** ask for supervisor, method, contribution, finding/status, authorship, venue, and funding details when material. **Do not:** fill gaps from field conventions. |

### 9.3 Transformation rules

| Output unit | Corporate transformation | Academic transformation |
|---|---|---|
| Entry focus | Relevant capability and result | Research preparation, contribution, and trajectory |
| Safe structure | **Action + artifact/problem + scope/constraint + verified outcome + relevant technology** | **Research question/problem + method/theory + data/system + individual contribution + result/status/output + supervisor/collaboration where useful** |
| Metrics | Prefer verified operational/business/quality measures; a precise qualitative outcome is better than an invented number | Prefer methodological scale, evaluation evidence, finding/status, output, and scholarly significance; preserve uncertainty |
| Technology | Select only job-relevant, evidenced stack | Select methods/tools needed to understand the work; prioritize theory/method over stack inventory |
| Team outcome | Attribute candidate contribution before team/company outcome | Attribute applicant contribution before lab/paper/project outcome |
| Length | Usually one concise bullet per distinct contribution; 2–5 bullets per recent relevant role | One or more concise descriptive bullets/lines per research role; complete publication citations in dedicated sections |

**Corporate example transformation pattern (not candidate content):**

`Implemented [specific component] for [system/context] using [verified technologies], supporting [verified scale/constraint] and resulting in [verified outcome].`

**Academic example transformation pattern (not candidate content):**

`Investigated [research question] using [method/theory] on [dataset/system]; contributed [specific individual work], yielding [verified result/status/output] under [supervisor/collaboration, if useful].`

## 10. Generator control logic and final quality gates

### 10.1 Mandatory decision sequence

1. Identify the requested track: corporate software engineering, industrial research, MSc/PhD admission, academic job, fellowship, or named funding scheme.
2. Load and timestamp the explicit posting/program/funder instructions, including document format, length/template, eligibility, and AI-use/disclosure rules.
3. Build the candidate evidence ledger; mark every fact Verified, User-asserted, Derived, Inferred, Unknown, or Conflicted.
4. Build the appropriate corporate requirement matrix or academic alignment matrix.
5. Select only relevant supported facts; never convert absence into a claim.
6. Generate with canonical labels, audience-specific tone, and format rules.
7. Validate the internal record against the companion JSON Schema and run semantic cross-reference/source checks that JSON Schema cannot perform.
8. Run all gates below; a failed hard gate blocks submission-ready output.

### 10.2 Quality gates

| Gate | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) | Severity |
|---|---|---|---|
| Truth gate | **Do:** map every factual clause and number to evidence. **Do not:** allow inferred facts. | **Do:** map every status, authorship, method, finding, grant, role, and number to evidence. **Do not:** allow scholarly inference to masquerade as fact. | Hard block |
| Conflict gate | **Do:** resolve title/date/metric/technology conflicts. **Do not:** choose the more favorable source silently. | **Do:** resolve publication/author/date/degree/funding/supervisor conflicts. **Do not:** export ambiguity. | Hard block |
| Instruction gate | **Do:** meet file, page, naming, and application requirements. **Do not:** apply defaults over explicit instructions. | **Do:** meet portal/funder/discipline template rules exactly. **Do not:** improvise required forms. | Hard block |
| Status gate | **Do:** label prototype, deployment, certification, and employment status correctly. **Do not:** upgrade status. | **Do:** label publication, funding, teaching, degree, and research status correctly. **Do not:** upgrade status. | Hard block |
| Parsing gate | **Do:** confirm readable text, linear order, canonical headings, and correctly extracted contact/employment/education. **Do not:** rely only on appearance. | **Do:** confirm searchable text and intact symbols/citations; parsing is secondary to readable rendering. **Do not:** ignore broken text. | Hard block for corporate; hard block for unreadable academic PDF |
| Alignment gate | **Do:** cover verified must-haves and prioritize evidence relevant to the role. **Do not:** pad with unrelated terms. | **Do:** demonstrate question/method/domain/output/trajectory fit. **Do not:** name-drop without a substantive connection. | Quality block |
| Human-scan gate | **Do:** make role identity, strongest relevant evidence, and contact data obvious on page one. **Do not:** lead with low-value detail. | **Do:** make current stage, field, strongest research evidence, and trajectory obvious early. **Do not:** bury defining work. | Quality block |
| Tone gate | **Do:** remove banned phrases, hype, vague adjectives, and repetitive AI cadence. **Do not:** substitute thesaurus verbs. | **Do:** remove hype and enforce epistemic precision. **Do not:** erase necessary disciplinary nuance. | Quality block |
| Integrity gate | **Do:** confirm compliant AI use, no hidden/stuffed/adversarial content, no identity-signal manipulation, and candidate ability to defend every claim. **Do not:** optimize for detector evasion. | **Do:** confirm programme AI policy/disclosure, attribution, research integrity, and applicant ability to defend every claim. **Do not:** use AI to replace the applicant's ideas, findings, or authorship. | Hard block |
| Consistency gate | **Do:** compare resume, form answers, LinkedIn/portfolio, and prior verified records. **Do not:** permit factual drift between variants. | **Do:** compare CV, statement, portal, transcript, ORCID/publication records, and referee data. **Do not:** permit status or trajectory contradictions. | Hard block |
| Visual gate | **Do:** inspect PDF/DOCX rendering for clipping, orphans, weak hierarchy, and density. **Do not:** deliver an unrendered source file. | **Do:** inspect every page, page numbers, bibliography wrapping, glyphs, and hyperlinks. **Do not:** assume compilation equals quality. | Quality block |
| Finalization gate | **Do:** remove notes, comments, metadata leaks, change tracking, and verification placeholders. **Do not:** submit a working draft. | **Do:** remove the same artifacts and confirm confidential/embargoed material is absent. **Do not:** expose review notes. | Hard block |

### 10.3 Machine-enforceable guardrail contract

The companion [ATS_Audience_Guardrails.schema.json](ATS_Audience_Guardrails.schema.json) is a JSON Schema Draft 2020-12 contract for the future generator's internal ledger. It validates target-instruction snapshots, sources, claims, project/publication records, selected output claims, AI/integrity controls, and all quality gates. It was compiled in strict Draft 2020-12 mode and tested with both a valid submission-ready record and a deliberately invalid record; the invalid record was rejected for inferred output, unresolved questions, placeholders, hidden text, unchecked AI policy, and failed gates.

Standard JSON Schema cannot prove that every string ID reference points to an existing object or that IDs are globally unique; those are semantic checks. The schema therefore requires explicit audit assertions (`claim_ids_unique`, `source_ids_unique`, `references_resolved`), and the future implementation must independently compute those assertions rather than trusting model output.

| ID | Schema insight | Corporate Rule (Do / Do Not) | Academic Rule (Do / Do Not) |
|---|---|---|---|
| JS-01 | Submission-mode hard boundary | **Do:** allow submission-ready output only when every selected claim is Verified, eligible User-asserted, or transparently Derived; all gates pass or are legitimately not applicable; placeholders and unresolved questions are absent. **Do not:** serialize Inferred, Unknown, Conflicted, confidential, or restricted claims into the resume. | **Do:** apply the same boundary to status, authorship, method, finding, degree, teaching, and funding claims. **Do not:** allow an unresolved scholarly claim into a final CV or funder form. |
| JS-02 | High-risk user assertions | **Do:** independently verify high-risk metrics, eligibility, security/compliance, production scale, or credentials before selection. **Do not:** mark a high-risk self-report submission-eligible merely because the user supplied it. | **Do:** independently verify high-risk authorship, publication acceptance, grant status/amount, finding, ethics approval, degree, and award claims. **Do not:** treat self-report as sufficient for these claims. |
| JS-03 | Metric completeness | **Do:** require measurement, unit, scope, time window, and defensible attribution context for every selected numeric claim. **Do not:** accept a bare percentage or large number. | **Do:** require the relevant dataset/sample, measure/unit, scope, time window or study context, and status. **Do not:** output decontextualized accuracy, citation, funding, or teaching counts. |
| JS-04 | Derived claims | **Do:** record input claim IDs, the calculation rule, and unit preservation. **Do not:** sum overlapping employment or convert estimates into measured impact. | **Do:** use only complete verified inputs for counts/durations and preserve academic definitions. **Do not:** derive independence, quality, novelty, or causal significance from counts. |
| JS-05 | Integrity controls | **Do:** require explicit false values for hidden text, keyword stuffing, adversarial encoding, identity-signal manipulation, and unsupported “humanizer” passes. **Do not:** make evasive tactics configurable optimizations. | **Do:** enforce the same controls and complete any required AI-use disclosure. **Do not:** weaken research-integrity or accessibility rules to influence detectors. |
| JS-06 | Semantic validation outside JSON Schema | **Do:** programmatically verify reference resolution, ID uniqueness, claim-to-sentence coverage, exact source content, extraction order, and rendered output. **Do not:** equate JSON-schema validity with factual truth. | **Do:** additionally compare publications, ORCID, transcripts, portal fields, funder form versions, and references. **Do not:** equate structural validity with academic compliance or scholarly accuracy. |

## 11. Compact system-prompt-ready guardrails

The following statements can be copied almost verbatim into the future generator prompt:

1. Never invent or estimate a metric, skill, technology, title, employer, date, role, publication, finding, authorship position, award, grant, teaching responsibility, certification, or link.
2. Treat missing information as **Unknown**, not false, zero, absent, current, or completed.
3. Use inferred facts only to formulate clarification questions; never place them in candidate-facing content.
4. Resolve source conflicts before generating submission-ready content; never choose the more impressive claim by default.
5. Preserve official titles and statuses. Functional clarifiers may be added only when factual and visibly secondary.
6. Never upgrade project stage, deployment stage, publication status, funding status, degree status, or teaching authority.
7. Never infer production readiness, security, scalability, compliance, novelty, causality, robustness, or research validity from intent or technology choice.
8. For corporate output, optimize verified requirement coverage and evidence—not keyword density. Use a single-column, parser-safe 1–2 page resume with canonical section labels.
9. For academic output, optimize research coherence, method/domain fit, contribution clarity, output status, trajectory, and reader navigation—not ATS density. Use a comprehensive CV unless the application mandates a specific format.
10. Explicit job/program/funder instructions override all formatting and ordering defaults, but never override truth.
11. Use plain, specific verbs. Ban generic AI phrasing, hype, self-praise, and unsupported superlatives.
12. Extract a provenance-backed fact ledger from websites and repositories before writing; do not summarize marketing copy directly into claims.
13. A repository dependency does not prove proficiency; a public repository does not prove open-source licensing; a live demo does not prove production deployment; commit count does not prove leadership.
14. In final mode, omit unresolved content. Never expose placeholders or verification notes in the submitted document.
15. Validate both machine extraction and human rendering, then cross-check all application materials for consistency.
16. Never claim or optimize for a universal ATS or AI-writing detector. Parsing, matching, application gates, fraud signals, identity checks, and human review are separate controls.
17. Never manipulate hidden text, encoding, device/contact/location signals, identity checks, or prose quality to evade screening; never add errors to appear human.
18. Treat job, programme, and funder instructions as versioned source evidence. Snapshot the live requirements and never reuse a prior-year limit without verification.
19. In submission-ready mode, every selected claim must be provenance-backed and submission-eligible, every required gate must pass, and unresolved questions/placeholders must be empty.
20. JSON Schema validates structure, not truth. Independently check source content, cross-references, global ID uniqueness, rendered output, and cross-document consistency.

## 12. Research basis and important corrections

- [Workday Resume Parsing](https://doc.workday.com/admin-guide/en-us/human-capital-management/recruiting/candidates/set-up-prospects-and-candidates/hdc1552497830785.html): parsing populates fields; results vary by format and word order; image-based styles are discouraged.
- [Workday Candidate Skills Match](https://doc.workday.com/admin-guide/en-us/human-capital-management/recruiting/candidates/candidate-skills-match/bmj1604095304483.html): required skills receive greater weight; this feature’s match calculation does not consider skill recency, duration, or total work experience.
- [Workday Recruitment Privacy Statement](https://www.workday.com/en-us/privacy/recruiting-privacy-statement.html) (June 2026): describes Workday's own use of extracted application/resume fields and suggested qualification grades. This is evidence of a possible configured workflow, not a universal customer setting.
- [Greenhouse Parsing Failures](https://support.greenhouse.io/hc/en-us/articles/200989175-Unsuccessful-resume-parse): documents specific failure risks including columns, tables, headers/footers, text boxes, graphics, and image files.
- [Greenhouse Boolean Search](https://support.greenhouse.io/hc/en-us/articles/202360199-Search-candidates-using-Boolean-queries): recruiters can use exact phrases, AND/OR/NOT, grouping, and wildcards.
- [Greenhouse Talent Matching](https://support.greenhouse.io/hc/en-us/articles/41396009937307-Talent-Matching) (June 2026) and [Talent Matching FAQ](https://support.greenhouse.io/hc/en-us/articles/41131616864283-Talent-Matching-Data-Processing-FAQ): employers define and weight calibration criteria; matching extracts a bounded set of resume fields and is assistive rather than an automatic advance/reject decision.
- [Greenhouse Auto-Reject](https://support.greenhouse.io/hc/en-us/articles/360000653472-Auto-reject): structured application answers can be configured as automatic gates, separate from resume matching.
- [Greenhouse Fraud Detection Security/Privacy FAQ](https://support.greenhouse.io/hc/en-us/articles/45397259312027-Fraud-Detection-and-Spam-Blocklist-Security-Privacy-FAQ) and [Fraud Policy Guide](https://support.greenhouse.io/hc/en-us/articles/44681941657243-Operational-readiness-guide-Fraud-Detection-policy) (2026): document device/contact risk signals, separate spam controls, recruiter review, and the fact that the fraud-report feature is not generative AI or an automated rejection decision.
- [Lever Resume Parsing](https://help.lever.co/s/article/Understanding-Resume-Parsing) and [Lever Candidate Search](https://help.lever.co/s/article/Searching-the-Database-for-Candidates): parsing populates profile data and recruiters can retrieve candidates through keywords/fields; behavior remains tenant/workflow dependent.
- [Lever 2026 Candidate-Fraud Guidance](https://www.lever.co/blog/what-the-rise-in-ai-powered-candidate-fraud-really-means-for-ta-teams): distinguishes acceptable assistive AI use from misrepresentation and describes layered identity, skill-validation, and fraud controls.
- [Oracle Taleo Candidate Management](https://docs.oracle.com/en/cloud/saas/taleo-enterprise/21b/otrec/candidate-management.html), [Application Flow Blocks](https://docs.oracle.com/en/cloud/saas/taleo-enterprise/otcug/r-applicationflowblocks.html), and [Getting Started](https://docs.oracle.com/en/cloud/saas/taleo-enterprise/20d/otrec/getting-started.html): document parsing, conceptual search, and screening as separate functions. These are publicly accessible legacy manuals, so release-specific limits or settings must not be presented as 2026 universal behavior.
- [iCIMS 2026 ATS Guide](https://www.icims.com/blog/what-to-look-for-in-an-applicant-tracking-system-complete-buyers-guide/), [AI Recruiting](https://www.icims.com/products/ai-recruiting-software/), and [Enterprise ATS](https://www.icims.com/products/hiring-software/enterprise-applicant-tracking-system/): document searchable parsing/prepopulation, search, comparison, ranking, matching, and configurable workflows.
- [Textkernel Resume Parsing Workflow](https://developer.textkernel.com/tx-platform/v10/resume-parser/overview/getting-started/): documents conversion, usability validation, OCR, parsing, and output stages.
- [Textkernel Data Model](https://developer.textkernel.com/Parser/master/data_model/) and [Technical Specifications](https://developer.textkernel.com/tx-platform/v10/resume-parser/overview/specs/): structured JSON/XML fields, normalization, derived information, and supported sections.
- [Textkernel–Sovren acquisition](https://www.textkernel.com/learn-support/blog/textkernel-acquires-sovren-to-become-the-global-leader-in-ai-powered-recruitment-technology/): Sovren joined Textkernel, explaining why current Tx Platform documentation is used for the Sovren-style parsing and matching discussion.
- [Textkernel Search & Match FAQ](https://developer.textkernel.com/tx-platform/v10/faq/): distinguishes human-authored search, automated match, bidirectional scoring, taxonomy normalization, and category weighting.
- [DaXtra Parser](https://www.daxtra.com/products/resume-parsing-software/): describes structured extraction across 150+ fields and skills taxonomies.
- [Harvard Resume Guidance](https://careerservices.fas.harvard.edu/resources/hes-create-impactful-resumes-and-cover-letters/) and [Penn Master’s Application Guidance](https://careerservices.upenn.edu/internship-and-job-applications-for-masters-students/): role tailoring, fact-based impact, canonical headings, simple one-column formatting, and cautious use of AI.
- [Princeton Graduate CV Guide](https://careerdevelopment.princeton.edu/guides/resume-cv-cover-letter-perspective-statement/cv-writing-guide-for-graduate-students), [Yale Academic CV Guide](https://ocs.yale.edu/%F0%9F%93%84-your-academic-cv/), [Oxford Academic CV Guidance](https://www.careers.ox.ac.uk/cvs), and [Cambridge PhD/Postdoc CV Guide](https://www.careers.cam.ac.uk/files/phdpostdoccvbook.pdf): a scholarly CV is comprehensive but audience-tailored, with research, teaching, outputs, awards/funding, and service ordered according to the opportunity.
- [Cornell CV Guidance](https://gradschool.cornell.edu/career-and-professional-development/pathways-to-success/prepare-for-your-career/take-action/resumes-and-cvs/): a CV is comprehensive, can span several pages, and emphasizes academic/research history.
- [Berkeley Academic CV Guidance](https://career.berkeley.edu/grad-students-postdocs/academic-job-search/the-cv-part-2-elements/): discipline-aware publication formatting, author visibility, and strict separation of accepted/forthcoming from submitted work.
- [MIT Graduate Statement Guidance](https://mitcommlab.mit.edu/eecs/commkit/graduate-school-statement-of-purpose/) and [Cornell Academic Statement Guidance](https://gradschool.cornell.edu/inclusion/recruitment/prospective-students/writing-your-statement-of-purpose/): committees seek concrete research preparation, accomplishment, goals, and program/lab fit.
- [NIH 2026 Common-Form Enforcement](https://grants.nih.gov/grants/guide/notice-files/NOT-OD-26-079.html), [NSF Senior/Key Personnel Documents](https://www.nsf.gov/funding/senior-personnel-documents), [ERC 2026 call guidance](https://erc.europa.eu/news-events/events/erc-grants-what-expect-2026-calls), [UKRI R4RI guidance](https://www.ukri.org/apply-for-funding/develop-your-application/resume-for-research-and-innovation-r4ri-guidance/), and [DFG CV FAQ](https://www.dfg.de/en/research-funding/proposal-funding-process/faq/cv): funding applications may require certified common forms, bounded track records, mandatory templates, or narrative evidence rather than a generic exhaustive CV.
- [NIST Text-to-Text Evaluation](https://www.nist.gov/publications/2024-nist-genai-pilot-study-text-text-evaluation-overview-and-results): detector performance is system-dependent. It does not establish a universal ATS résumé-authorship detector.

### Corrections to common ATS myths

1. **Myth:** Every ATS assigns one hidden universal score. **Correction:** parsing, search, matching, knockout rules, third-party screening, and human review are distinct and configurable.
2. **Myth:** Keyword density determines selection. **Correction:** exact terms can support retrieval, but field mapping, required-skill weighting, synonyms, evidence, eligibility questions, and human judgment vary by system.
3. **Myth:** An ATS automatically rejects any resume it cannot parse. **Correction:** workflows vary; Greenhouse, for example, can retain an attachment after an unsuccessful parse, while separate application rules can auto-reject based on answers.
4. **Myth:** PDF is always best. **Correction:** text-based PDF is usually safe, but the posting/portal’s requested format controls, and DOCX is commonly supported.
5. **Myth:** More keywords are always better. **Correction:** unsupported or repeated terms reduce credibility and can misrepresent the candidate; one truthful exact term plus contextual evidence is stronger.
6. **Myth:** Academic CVs should be ATS-optimized like resumes. **Correction:** the uploaded CV is primarily a human-readable scholarly record, while structured portal fields and application documents may be evaluated separately.
7. **Myth:** Funding applications always accept a long CV. **Correction:** major funders increasingly use scheme-specific common forms, constrained track records, or narrative CVs.
8. **Myth:** Mainstream ATS products universally detect and reject AI-written resumes. **Correction:** current public documentation supports matching, application gates, fraud/device signals, identity verification, and human review; it does not establish one reliable cross-vendor résumé-authorship detector.
9. **Myth:** Adding typos or “humanizing” prose prevents AI detection. **Correction:** deliberate errors, paraphrase noise, and encoding tricks damage credibility/accessibility and can introduce factual drift; provenance and defensible specificity are the safe controls.
10. **Myth:** A valid JSON guardrail record proves the résumé is true. **Correction:** schema validation proves structural compliance only; evidence resolution, source content, cross-document consistency, and rendering require independent checks.

---

**End state for Step 1:** the future System Prompt should implement two independent generation branches sharing one truth/provenance engine, with audience-specific selection, labels, tone, formatting, and QA gates.
