CV parsing is the automated extraction of structured data—like name, contact details, employers, and skills—from a CV.
What is CV parsing and why does it matter?
CV parsing turns an unstructured document into data that an ATS or database can read. Recruitment agencies rely on it to speed up candidate screening and reduce manual data entry. But parsing isn't perfect. Formatting quirks, tables, or unusual layouts often cause errors that waste time rather than save it.
How CV parsing works
CV parsers use algorithms to scan the text and identify key fields. Most rely on pattern recognition and keyword matching to locate:
- Candidate name and contact details
- Employment history and job titles
- Education and qualifications
- Skills and certifications
Parsers break the CV into sections, then extract and tag the information. More advanced parsers use natural language processing to handle variations in phrasing. Still, they expect CVs to follow common formatting conventions.
Parsing accuracy depends on clear structure. Parsers struggle when:
- Text is embedded in tables, text boxes, or columns
- Fonts or special characters confuse extraction
- CVs contain images, logos, or graphics
- Layouts vary widely from the parser's training data
Errors often show as missing or scrambled data, fields assigned to the wrong category, or duplication. This leads to extra manual correction, defeating the purpose.
Examples of parsing failures
- Bullhorn scrambles CVs with tables and columns, mixing up job titles and dates. The fix—remove tables and use simple, linear layouts.
- Some parsers read headers or footers as current employer details, inflating the candidate's work history.
- PDFs saved as images or scanned documents provide no text to extract, so the parser returns blank fields.
- Multiple email addresses or phone numbers in the CV confuse parsers about the correct contact.
Understanding these pitfalls helps recruiters prepare CVs that parse more reliably.
Related terms
- Applicant Tracking System (ATS): The software that receives parsed CV data to manage candidate workflows. See more in our ATS guides.
- Data anonymisation: The process of removing personal identifiers from CVs before submission, important for GDPR compliance. Our compliance tools cover this in depth.
- Structured data: Information organised into defined fields like name, dates, and skills, as opposed to free text in a CV. Parsing creates structured data from unstructured documents.
- CV formatting: How the CV is laid out affects parsing success. Simple, consistent formatting improves extraction accuracy.
Where Distill fits
Distill addresses common parsing errors by cleaning and standardising CVs before submission. It strips out problematic elements like photos, personal contact details, and graduation years that can cause compliance issues or parsing failures. Distill also reformats CVs to match client ATS specifications, reducing manual fixes.
If you send 20+ CVs a week to Bullhorn or similar ATS users, Distill can automatically apply these fixes to save time and reduce errors.
Try Distill free to see how it improves parsing accuracy and compliance for your agency.