HireLayer logoHireLayer

How to Parse a Resume PDF into JSON with Python or Node.js

Compare local PDF extraction, OCR, and a resume parser API, then follow working Python and Node.js examples for turning a resume PDF into structured JSON.

Published · 9 minutes read

A resume PDF being converted into structured JSON fields

A PDF library can extract characters from some resumes, but it does not automatically know which words are a candidate’s name, which dates belong to a role, or where one education entry ends. Converting a resume PDF to useful JSON therefore has two jobs: read the document and map its content into fields your application understands.

What “PDF to JSON” actually involves

A PDF may contain selectable text, or it may be a scan made of page images. In the first case, a library such as pypdf can retrieve text. In the second, text extraction alone has nothing to read; an OCR step must first recognize characters in the image. Neither step by itself produces a reliable candidate object.

A useful result is a schema such as info_candidate, work_experiences[], educations[], and skills[]. The integration then maps those fields into your own database and keeps missing or uncertain values reviewable. Avoid treating an empty field as proof that the candidate lacks that information.

Three implementation approaches

ApproachUseful whenTrade-off
PDF text extraction + rulesYou need text from a controlled document set.You own layout variation, OCR, and field mapping.
OCR + custom extractionScanned files are common and you need control.More components to tune, monitor, and maintain.
Resume parser APIYou need candidate fields without building the extraction pipeline.You depend on a vendor contract, limits, and pricing.

Start with the simplest approach that meets your document coverage and quality requirements. If you choose an API, compare its accepted formats, request limits, response schema, failure behavior, and data handling before committing to the integration.

The examples below call the HireLayer CV Extract resume parsing API, which takes one file per multipart request and returns candidate details, work history, education, languages and skills as JSON.

Parse a PDF resume with Python

This server-side example sends one PDF to the V3 endpoint and reads two fields from the JSON response. Install the requests package and set the API key in the server environment. Never put a production secret in browser code or a public repository.

import os
import requests

api_key = os.environ["HIRELAYER_API_KEY"]
url = "https://onlineresumeparser.com/api/v3/parser"

with open("resume.pdf", "rb") as resume:
    response = requests.post(
        url,
        headers={"X-API-Key": api_key},
        files={"file": ("resume.pdf", resume, "application/pdf")},
        data={
            "application_id": "candidate-123",
            "do_not_store_data": "true",
        },
        timeout=150,
    )

response.raise_for_status()
parsed = response.json()

candidate_name = parsed.get("info_candidate", {}).get("full_name")
experience = parsed.get("work_experiences", [])
print({"has_candidate_name": bool(candidate_name), "experience_count": len(experience)})

The request uses multipart/form-data, sends the file as the required file field, and puts credentials in X-API-Key. The API accepts an optional application_id reference. Set do_not_store_data to true when the parsed data must not be retained; decide separately how your own service stores and deletes uploaded documents and results.

Parse a PDF resume with Node.js

For this example, install axios and form-data. Run it in a backend worker or API route so the key never reaches the candidate’s browser.

import fs from 'node:fs'
import axios from 'axios'
import FormData from 'form-data'

const form = new FormData()
form.append('file', fs.createReadStream('resume.pdf'))
form.append('application_id', 'candidate-123')
form.append('do_not_store_data', 'true')

const response = await axios.post(
  'https://onlineresumeparser.com/api/v3/parser',
  form,
  {
    headers: {
      ...form.getHeaders(),
      'X-API-Key': process.env.HIRELAYER_API_KEY,
    },
    timeout: 150_000,
  }
)

const parsed = response.data
const candidateName = parsed.info_candidate?.full_name ?? null
console.log({
  hasCandidateName: Boolean(candidateName),
  experienceCount: parsed.work_experiences?.length ?? 0,
})

The returned object is the parser response, not a guarantee that every field is populated on every resume. Check the documented response and error codes, then validate the fields that your workflow depends on. A successful HTTP response can still contain warnings or partial upstream results; preserve that context for downstream review.

Try the same PDF in the live demo

Compare the structured response with the original document, then use the API reference to check the exact request and response contract.

Map parser output into your schema

Keep a small adapter between a vendor response and your internal model. For example, store the parser’s request_id for support and traceability, associate the result with your own candidate identifier, and map each work-history entry into the fields your ATS requires. Preserve the source response only if your retention policy allows it.

{
  "candidate": {
    "name": "Alex Morgan",
    "email": "[email protected]"
  },
  "experience": [
    {
      "employer": "Northstar Labs",
      "title": "Software Engineer",
      "startDate": "2022-03-01"
    }
  ]
}

This is a shortened illustrative shape for an application’s own schema, not a captured API response. The live API response uses fields such as info_candidate.full_name and work_experiences[].company_name; see the V3 API documentation for its current structure.

If this parser will feed a recruiting platform, see the ATS integration guide. For processing a folder of resumes, the bulk pipeline guide covers queues and retries.

Test difficult PDFs before launch

  • Selectable text: verify names, email addresses, and dates in ordinary exports.
  • Image-only scan: check OCR-dependent results and identify unreadable uploads.
  • Two-column layout: inspect reading order and whether roles are associated with the right dates.
  • Unusual chronology: include overlapping roles, missing end dates, and education mixed with experience.
  • Format and size: test only files your upload flow permits; the documented encoded request limit is 6 MiB.

Use consented, synthetic, or properly anonymized test files. Record field-level errors against a human-checked reference set; a single end-to-end “looks right” check hides whether names, dates, or employers fail more often. Keep the test corpus and its limitations with the result.

Frequently asked questions

Can a Python PDF library return candidate JSON directly?

No. A PDF text library extracts text. You still need OCR for image-only pages and logic that identifies and maps candidate fields.

Does the API accept one file or a whole batch?

The documented V3 parser request accepts one file. For multiple resumes, submit separate requests from a controlled server-side queue.

What should I do with a missing field?

Represent it as unknown or empty according to your schema, preserve any warnings, and send consequential fields for human review.

Sources and further reading

  1. HireLayer API documentation: Parse a resume (V3)
  2. pypdf documentation: Extract text from a PDF
  3. OWASP File Upload Cheat Sheet

Louis Desclous

Published on · Reading time: 9 minutes