USE CASE / AI TRAINING DATA

Real job postings
for training data.

Millions of deduplicated job postings with full descriptions, titles, pay and locations to train, fine-tune and evaluate models.

10.1Mjob postings
200k+company job boards
24hevery board re-checked
500jobs a month free

Why Headcount Labs

Full text

Every job carries the employer’s full description, as plain text for training or HTML when you need the original layout.

No repeats

Reposted and syndicated copies are merged into one record, so duplicates don’t skew your training set or leak between splits.

Structured fields

Titles, employers, locations, work arrangement, employment type and posted pay come as separate fields, ready to use as labels.

Bulk pulls

Page through any filter with a cursor, up to 250 jobs per REST call, and save straight to JSONL. Pay only for the jobs you take.

How it works

Ask your agent in plain words. It calls Headcount Labs through MCP, or your code calls the same filters over the REST API.

Claude Code · headcountlabs connected
Build a training set of 50,000 US software engineering job posts with pay and full descriptions, saved as JSONL.
count_jobs{ filters: { title: "software engineer", country: "US" } }
search_jobs{ filters: { title: "software engineer", country: "US", description_format: "text" }, limit: 25, cursor: "…" }
Saved 50,000 deduplicated postings to data/swe_jobs.jsonl, skipping any without posted pay or a full description. Each row has the title, employer, location, pay and description in plain text.

Example conversation.

The same request over REST

curl -G https://api.headcountlabs.com/v1/jobs \
  --data-urlencode "title=software engineer" \
  --data-urlencode "country=US" \
  --data-urlencode "description_format=text" \
  --data-urlencode "limit=100" \
  -H "Authorization: Bearer $HEADCOUNTLABS_KEY"

What you get

  • Full description as plain text or HTML
  • Job title and employer
  • City, state and country
  • Pay, when the employer posts it
  • Work arrangement and employment type
  • Posted date and source links for every record

More use cases

Start with 500 jobs free.

No card needed. Connect the API or MCP in minutes.

Try for free