A PDF library can extract characters from some resumes, but it does not automatically know which words are a candidate’s name, which dates belong to a role, or where one education entry ends. Converting a resume PDF to useful JSON therefore has two jobs: read the document and map its content into fields your application understands.
What “PDF to JSON” actually involves
A PDF may contain selectable text, or it may be a scan made of page images. In the first case, a library such as pypdf can retrieve text. In the second, text extraction alone has nothing to read; an OCR step must first recognize characters in the image. Neither step by itself produces a reliable candidate object.
A useful result is a schema such as info_candidate,
work_experiences[], educations[], and
skills[]. The integration then maps those fields into your
own database and keeps missing or uncertain values reviewable. Avoid
treating an empty field as proof that the candidate lacks that
information.
Three implementation approaches
| Approach | Useful when | Trade-off |
|---|---|---|
| PDF text extraction + rules | You need text from a controlled document set. | You own layout variation, OCR, and field mapping. |
| OCR + custom extraction | Scanned files are common and you need control. | More components to tune, monitor, and maintain. |
| Resume parser API | You need candidate fields without building the extraction pipeline. | You depend on a vendor contract, limits, and pricing. |
Start with the simplest approach that meets your document coverage and quality requirements. If you choose an API, compare its accepted formats, request limits, response schema, failure behavior, and data handling before committing to the integration.
The examples below call the HireLayer CV Extract resume parsing API, which takes one file per multipart request and returns candidate details, work history, education, languages and skills as JSON.
Parse a PDF resume with Python
This server-side example sends one PDF to the V3 endpoint and reads two
fields from the JSON response. Install the requests package
and set the API key in the server environment. Never put a production
secret in browser code or a public repository.
import os
import requests
api_key = os.environ["HIRELAYER_API_KEY"]
url = "https://onlineresumeparser.com/api/v3/parser"
with open("resume.pdf", "rb") as resume:
response = requests.post(
url,
headers={"X-API-Key": api_key},
files={"file": ("resume.pdf", resume, "application/pdf")},
data={
"application_id": "candidate-123",
"do_not_store_data": "true",
},
timeout=150,
)
response.raise_for_status()
parsed = response.json()
candidate_name = parsed.get("info_candidate", {}).get("full_name")
experience = parsed.get("work_experiences", [])
print({"has_candidate_name": bool(candidate_name), "experience_count": len(experience)})
The request uses multipart/form-data, sends the file as the
required file field, and puts credentials in
X-API-Key. The API accepts an optional
application_id reference. Set
do_not_store_data to true when the parsed data
must not be retained; decide separately how your own service stores and
deletes uploaded documents and results.
Parse a PDF resume with Node.js
For this example, install axios and form-data.
Run it in a backend worker or API route so the key never reaches the
candidate’s browser.
import fs from 'node:fs'
import axios from 'axios'
import FormData from 'form-data'
const form = new FormData()
form.append('file', fs.createReadStream('resume.pdf'))
form.append('application_id', 'candidate-123')
form.append('do_not_store_data', 'true')
const response = await axios.post(
'https://onlineresumeparser.com/api/v3/parser',
form,
{
headers: {
...form.getHeaders(),
'X-API-Key': process.env.HIRELAYER_API_KEY,
},
timeout: 150_000,
}
)
const parsed = response.data
const candidateName = parsed.info_candidate?.full_name ?? null
console.log({
hasCandidateName: Boolean(candidateName),
experienceCount: parsed.work_experiences?.length ?? 0,
})
The returned object is the parser response, not a guarantee that every field is populated on every resume. Check the documented response and error codes, then validate the fields that your workflow depends on. A successful HTTP response can still contain warnings or partial upstream results; preserve that context for downstream review.
Try the same PDF in the live demo
Compare the structured response with the original document, then use the API reference to check the exact request and response contract.
Map parser output into your schema
Keep a small adapter between a vendor response and your internal model.
For example, store the parser’s request_id for support and
traceability, associate the result with your own candidate identifier,
and map each work-history entry into the fields your ATS requires.
Preserve the source response only if your retention policy allows it.
{
"candidate": {
"name": "Alex Morgan",
"email": "[email protected]"
},
"experience": [
{
"employer": "Northstar Labs",
"title": "Software Engineer",
"startDate": "2022-03-01"
}
]
}
This is a shortened illustrative shape for an application’s own schema, not a
captured API response. The live API response uses fields such as
info_candidate.full_name and
work_experiences[].company_name; see the V3 API documentation for its current structure.
If this parser will feed a recruiting platform, see the ATS integration guide. For processing a folder of resumes, the bulk pipeline guide covers queues and retries.
Test difficult PDFs before launch
- Selectable text: verify names, email addresses, and dates in ordinary exports.
- Image-only scan: check OCR-dependent results and identify unreadable uploads.
- Two-column layout: inspect reading order and whether roles are associated with the right dates.
- Unusual chronology: include overlapping roles, missing end dates, and education mixed with experience.
- Format and size: test only files your upload flow permits; the documented encoded request limit is 6 MiB.
Use consented, synthetic, or properly anonymized test files. Record field-level errors against a human-checked reference set; a single end-to-end “looks right” check hides whether names, dates, or employers fail more often. Keep the test corpus and its limitations with the result.
Frequently asked questions
Can a Python PDF library return candidate JSON directly?
No. A PDF text library extracts text. You still need OCR for image-only pages and logic that identifies and maps candidate fields.
Does the API accept one file or a whole batch?
The documented V3 parser request accepts one file. For multiple resumes, submit separate requests from a controlled server-side queue.
What should I do with a missing field?
Represent it as unknown or empty according to your schema, preserve any warnings, and send consequential fields for human review.
Sources and further reading
- HireLayer API documentation: Parse a resume (V3)
- pypdf documentation: Extract text from a PDF
- OWASP File Upload Cheat Sheet
Louis Desclous
Published on · Reading time: 9 minutes

