Guide · Preparing your document

Preparing your document.

What converts well, what to skip, every format explained, ownership rules, and a 7-item checklist before you upload. Read this once and you'll never have to wonder why a build came back wrong.

§1 — What converts beautifully

These source types consistently produce sharp skills.

SkillPill's pipeline extracts knowledge from your document, plans a structured skill around its concepts, synthesises the chapters, and then runs an automated QA pass that traces every claim back to a specific passage in your source. That last step is the key one: it means the quality of the skill reflects the quality of the source. Structured, sectioned, well-organised documents produce skills that answer confidently and cite precisely. Loose, fragmented ones produce skills that are proportionally vaguer.

The following source types have a strong track record:

What works well
  • + Structured non-fiction books, 200–600 pagesChapters become on-demand modules. Ideal length — enough depth to get genuine retrieval without hitting the Trial page cap.
  • + Brand guides and style guidesRich in defined terms, do/don't rules, and repeatable patterns — exactly the shape the pipeline was designed for.
  • + Methodology docs and SOPsStep-by-step process structure converts cleanly into cheatsheet and patterns files.
  • + Course material and curriculaModule-by-module layout maps naturally to the chapters/ folder structure.
  • + Research vaults and literature reviewsDense citation structure means the QA pass has a lot to anchor to. Very high pass rates.
Avoid for your first skill
  • Scanned PDFs without a text layerThe pipeline needs selectable text. Image-only pages extract to nothing.
  • Password-protected filesThe pipeline cannot open encrypted documents. The job will fail and credits refund automatically.
  • Pure image decks (PowerPoint / Keynote with no speaker notes)Slide visuals don't extract. Text-heavy decks with thorough notes work fine.
  • Fiction and narrative writingSkills are for applied knowledge. A novel can be processed, but the result is a study tool, not a usable assistant — probably not what you want.
  • Files over your plan's page capTrial cap is 500 pages. Submit a portion of a longer work rather than the full file.
Why structure matters

The pipeline's planning step divides your source into topic clusters, then assigns each cluster to a chapter in the skill. A document with clear headings and sections — even simple H2/H3 structure — produces a skill with crisp chapters. A document that is a single wall of text still works, but the chapter boundaries will be less precise. If you have any control over the source before you upload, add section headings. Even a basic outline makes a measurable difference to the result.

The QA pass is the other reason structure matters. It works by taking each synthesised claim and finding the passage in your source that supports it. If the pipeline cannot find a supporting passage, the claim does not ship. Documents with clearly written, factual statements (rather than abstract implications or unstated assumptions) produce the highest QA pass rates — meaning more of your knowledge makes it into the final skill intact.

§2 — What to avoid

Problems to catch before you upload.

Scanned PDFs without a text layer

This is the most common reason builds fail. A scanned PDF is a photograph of a page — the text is not extractable because it is stored as image data, not characters. The QA pass has nothing to anchor to, so the pipeline fails and your credits are refunded automatically.

How to check whether your PDF has a text layer: Open the file in any PDF viewer and try to click and drag to select text. If you can highlight words, the file has a text layer and will work. If your cursor turns into a crosshair and you can only drag to draw a rectangle, the page is an image and needs OCR first.

To add a text layer to a scanned PDF, run it through OCR. Google Drive does this automatically if you open the PDF in Drive and then open with Google Docs — download the resulting Docs file as a PDF or DOCX. Adobe Acrobat has a built-in OCR step under Tools > Enhance Scans. Smallpdf and ilovepdf.com both offer free online OCR.

Password-protected files

If your file requires a password to open, the pipeline cannot read it. Remove the password protection before uploading. In Acrobat: File > Properties > Security > change Security Method to No Security. In Word: Review > Protect Document > and clear the password. Then re-save and upload the unprotected version.

Files over the page cap for your plan

Trial accounts are capped at 500 pages. If your document is longer, upload the portion that matters most for the skill you want — the chapters covering your core method, for example, rather than the full unabridged text including appendices. You can always build additional skills from other portions later.

The pipeline page count is based on the PDF page count or the equivalent for other formats. A 600-page EPUB will be treated as roughly 600 pages for credit calculation purposes. See the credit estimator on the upload screen for the exact cost before you confirm.

Mixed-language documents

If your document contains sections in multiple languages, the pipeline processes the whole document using the dominant-language cost multiplier. Bilingual documents where roughly equal portions are in two languages will cost more to process because the token count is higher and the extraction step is more complex. If you have a bilingual version and a single-language version of the same document, use the single-language one.

§3 — File formats

Seven formats supported. Here is what to expect from each.

All seven formats below are accepted at upload. The differences between them come down to how cleanly the text extracts and how much structure survives the extraction step.

Format Extraction quality Notes
PDF Good if text-based Depends entirely on whether the file has a text layer. Text-based PDFs extract cleanly. Scanned image PDFs extract nothing. Use the text-selection test described in §2 before uploading.
EPUB Excellent EPUB files contain structured HTML internally. Chapter boundaries, headings, and semantic markup are preserved during extraction, which gives the planning step precise signals. If you have a choice between PDF and EPUB, choose EPUB.
DOCX Excellent Word documents with heading styles (Heading 1, Heading 2) produce the clearest chapter structure. If your Word file uses manual bold formatting instead of paragraph styles, the extraction still works but the chapter divisions will be less precise.
MOBI Good MOBI is the original Kindle format. The pipeline converts MOBI files internally before processing. Quality is close to EPUB because the underlying structure is similar. If your ebook reader app can export to EPUB, prefer that over MOBI.
MD Excellent Markdown files extract cleanly and the heading hierarchy (# through ####) maps directly to the planning step's chapter detection. If you maintain any documentation or notes in Markdown, it is an ideal source format.
HTML Good HTML files are stripped of markup and the body text is extracted. Semantic heading tags (<h1> through <h4>) are preserved as structural signals. Navigation menus, footers, and sidebar content are filtered out during extraction, so single-page HTML docs work better than multi-page sites.
RTF Good Rich Text Format is handled similarly to DOCX. Text and basic formatting extract cleanly. Tables and complex inline elements may lose some fidelity, but the body text and headings come through well.
Quick rule

If you have any choice in format, use EPUB or MD first, DOCX second, PDF third. EPUB and Markdown give the pipeline the most structural information to work with. PDF is fine when it is the only option, as long as you have confirmed the text layer is present.

§4 — Ownership & your IP

You must own or have rights to convert the material. Here is what that means in practice.

At the upload screen, you will see a checkbox: "I own this material and have the right to convert it." This is not a formality — it is the only gate between your document and the pipeline. We take it seriously because your ability to use the resulting skill for professional work depends on whether you actually have the rights to that material.

You own or have rights to convert a document if any of these apply:

  • You wrote it yourself — a book, a brand guide, a methodology document, an SOP, course material, research notes.
  • You are employed by or contracted to the organisation that produced it, and converting internal documents for internal AI use falls within your role.
  • You have a licence that explicitly permits derivative works or personal-use conversions — some professional frameworks and style guides are published under licences that allow this. Read the licence before ticking.
  • The content is in the public domain.

You should not tick the ownership checkbox for a document if: you bought a commercial ebook or physical book (ownership of a copy does not transfer the IP rights to convert it for re-use), you downloaded a PDF from a website without a licence permitting conversion, or the document belongs to a client whose contract does not give you IP rights to the content.

What SkillPill does with your file

Source deleted after conversion — typically within minutes, never longer than 7 days. We never train any AI on your content. The skill files that result from the build are stored encrypted in your account and accessible only to you. You can delete them at any time from your dashboard.

Your source file is used for one purpose: to generate the skill. The pipeline reads it, extracts text from it, and then deletes it. It is not retained as a training dataset, it is not accessible to other users, and it is not used to improve the models. The skill files that are produced — the SKILL.md, chapters, glossary, cheatsheet, and patterns files — are yours. You can edit them, share them, or delete them. They reflect your knowledge because they were built from your source.

If you delete your account, your skill files are deleted with it. Skills are not copied to a backup or shared with SkillPill staff unless you explicitly share them through the sharing feature (available on Plus and above).

§5 — One document, one skill

Scope advice: one coherent source beats a bundle of everything.

The pipeline works best when the source is a coherent, single-topic document. The planning step tries to identify the primary framework or method in the source and organise everything else around it. When a single PDF contains three unrelated topics — say, a financial methodology, a product roadmap, and a set of brand guidelines — the resulting skill is weaker than three separate skills built from clean, single-topic sources.

If you have a vault of related documents, the right approach is to split by playbook or topic and build one skill per topic:

  • One skill for your sales methodology — the playbook, MEDDIC scripts, objection handling, ICP criteria, qualification frameworks.
  • One skill for your brand voice — the style guide, do/don't vocabulary, CTA patterns, tone register by audience.
  • One skill for your research vault — the literature synthesis, key findings, your annotated notes on the most relevant sources.

This is not about credit cost — it is about precision. A skill that is about one thing gives your AI a clear context to operate in. When you invoke it, the AI knows what domain it is in and what vocabulary to use. A skill that tries to be about everything tends to give answers that drift between contexts.

Naming your skill well

When you confirm the build, you can give the skill a name. This name becomes the slug used to invoke it in Claude Code — so sdr-qualification-playbook becomes /sdr-qualification-playbook, and brand-voice-acme becomes /brand-voice-acme.

A good slug is short, specific, and contains the topic. You will type it often. Avoid generic names like my-skill or notes. If you have multiple skills from the same domain, namespace them clearly: sales-playbook-q2 versus sales-playbook-q3 is easier to distinguish in a session than two skills both named sales.

Updating a skill later

Skills are plain .md files. If your source evolves — a new edition of the book, an updated brand guide, a revised SOP — you can re-upload the updated document and build a new version of the skill. Use a version number in the skill name to keep track: v1.0 in the first build, v2.0 after a major update. Old skills stay in your account until you delete them, so you can compare the two.

§6 — Pre-upload checklist

Seven things to verify before you click Upload.

Run through this before every build. Most failed jobs trace back to something on this list.

Pre-upload checklist · 7 items
01
Text is selectable in the file. Open the file, click into the text, and try to drag-select a sentence. If it highlights, you have a text layer. If your cursor draws a box around a region instead of selecting characters, the file is an image scan and will fail extraction.
02
No password protection. The file should open without a password. If it requires one, remove the protection before uploading (see §2 for how).
03
Page count is within your plan's cap. Trial cap: 500 pages. If your file is longer, trim it to the relevant portion before uploading. The credit estimator on the upload screen shows the page count and cost before you commit.
04
The document is single-topic and coherent. One method, one voice guide, one playbook — not a mixed bag. A focused source produces a focused skill.
05
You own or have rights to convert this material. Re-read §4 if you are unsure. Ticking the ownership checkbox on a file you do not have rights to is a terms-of-service violation and creates legal exposure for you.
06
You have a name for the skill ready. Think of the slug you will type to invoke it. Short, specific, lowercase, hyphens instead of spaces. Example: brand-voice-acme-2026.
07
You have read the install guide. Building takes up to 24 hours. Use the wait to read the install guide so you are ready the moment your skill arrives. Link: guide/install.html

Once your skill arrives, the quality of the answers you get will directly reflect the quality of the source you uploaded. Take five minutes on this checklist and you will not have to re-run a build.