What the market is hiring for
Open data roles from six sources: four remote-first job boards, any company ATS board we choose to watch, and our own ASEAN donor/UN procurement feed — which is the only one of the six that reaches Cambodia and Lao PDR with any depth. Each card shows the role family we classified it into, the skills extracted from the full job description, and how many of Data ColabX's service pillars the posting touches.
| Role | Employer | Country | Family | Level | Skills | DNA fit | Posted |
|---|
The skills behind the titles
Counted from the full job description, not the title — which is the whole reason the pipeline stores the document. A posting called "Senior Data Scientist" tells you almost nothing; its JD tells you whether the work is dbt and Airflow, or PyTorch, or a Power BI refresh. Percentages are the share of currently-open roles demanding each skill.
Skill demand by group
Share of open roles mentioning each skill.
Role families
How the open roles split across the taxonomy.
Where the roles are
Open roles by region. The shape of this chart is a statement about source coverage as much as about the market — see the method note below.
Market demand against our own DNA
Each posting is scored against the eight service pillars that describe what Data ColabX is. Five are pillars we already sell — data science and analytics, data collection and pipelines, geospatial intelligence, digital solutions and MIS, statistics and M&E with safeguards. Three are the adjacent capability we are deliberately building toward. Green means the market is buying work we already do; rose means the market is buying work we are still building. The gap is the training plan.
Is Agentic AI actually in the job specs?
Plenty of people will tell you the data job is being rewritten by agents. This panel does not take their word for it — it counts. Every open posting is tagged when its description actually demands agent frameworks, LLM and RAG work, pipeline engineering, data collection or scraping, MLOps, or governance. The bars are today's share; the chart is our own daily series, which grows by one real observation a day.
Era signals in today's postings
Share of open roles whose description demands each capability.
Our own daily series
—
Ten years back, ten years forward
Before any chart: what we can actually stand behind. A ten-year history of data-science job postings cannot be scraped retrospectively, at any budget, and this page will not pretend otherwise. Job boards delete closed postings rather than archiving them. "Data Scientist" only became its own US occupation code in the 2018 revision, so even official statistics have no comparable series before 2019. And the BLS timeseries API returns data for the current reference year only — we checked, and 2019 through 2024 come back empty for this occupation. So here is the split.
What we hold, first-party
- —
- Every posting's full job description, stored and searchable — the evidence behind every skill count on this page.
- One aggregate snapshot per day, written before the feed rotates. Once a day closes its breakdown cannot be rebuilt, which is exactly why the series starts now and not later.
What cannot be reconstructed
- Postings that closed before we started watching. Boards remove them; there is no archive to scrape.
- A like-for-like official series before 2019 — the occupation code did not exist under this name.
- Any Cambodian statistic for data-science or data-engineering employment. We looked. That absence is itself a finding, and it is the argument for collecting our own.
What is a model, clearly labelled
- The forward curve below. It is arithmetic over published anchors with the assumptions on screen and adjustable — not a forecast anyone published.
- Dashed lines are always projections. Solid lines are always observations.
- Figures marked DERIVED are our arithmetic on someone else's published number; the arithmetic is stated in the citation.
Published anchors — every figure carries its source
Forward scenario — an explicit model, not a forecast
Anchored on the BLS base of 275,600 US data-science jobs in 2025 and their projected 35% growth to 2035. Move the sliders to see how sensitive the shape is to assumptions you may not share — that sensitivity is the point.
Read this as a sensitivity tool, not a prediction. The defensible claim underneath it is narrower and more useful: the work is shifting from producing analysis by hand toward specifying, governing and auditing systems that produce it — which is why pipelines, data collection and governance carry as much weight on this page as modelling does.
Cambodia: counting what nobody counts
We are based in Phnom Penh, so this is the section that has to be honest rather than flattering. There is no official Cambodian statistic for data-science employment — not a small one, none. The global job boards that fill most of this page effectively do not post Cambodia-based data roles at all. What reaches this market is donor and UN work: MIS builds, survey data collection, M&E systems, GIS mapping. That is the real shape of the Cambodian data job today, and it is why our own feed is the only forward series that exists for it.
Cambodia-based data roles open right now, in our feed.
Cambodia data roles we have recorded since tracking began — the number that grows from here.
ASEAN data roles recorded, for the regional comparison Cambodia needs to be read against.
What the Cambodian data job looks like
Skills Cambodian postings ask for
Published context
From data science to the foundations beneath it
A training plan ordered by what the feed actually demands, not by what is fashionable to teach. The demand figure on each rung is live — it is the share of currently-open roles asking for that rung's skills, so the ladder re-orders itself as the market moves. Note where the mathematics sits: underneath, not at the end. Machine learning taught without statistics produces people who can call a library and cannot defend a result.
How this page is built
Published in full so the numbers can be argued with. Every count on this page traces to a specific matched term in a specific stored document.
01 — Fetch
scripts/scrape-ds-jobs/ polls four remote-first board APIs, any company ATS
board listed in data/ds-job-boards.json (Breezy, Greenhouse and Lever are
supported — adding an employer is one line, no code change), and re-reads our own ASEAN
donor/UN feed. Sources run through allSettled, so one board failing costs
that board's rows and nothing else.
02 — Capture the document
Most boards and ATS vendors publish a schema.org/JobPosting block containing
the complete job description. lib/jobposting.js reads it, which is why any
job URL can be turned into a stored TOR without a per-vendor adapter. Attachment links
(TOR, SOW, RFP, annex PDFs) are detected and recorded alongside.
03 — Gate, classify, extract
lib/taxonomy.js decides whether a posting is a data role, which family it
belongs to, and which skills its description demands. Procurement notices get a separate
gate because they title a deliverable, not a post. The rules are plain regex —
inspectable, and covered by test-taxonomy.mjs.
04 — Store
Output lands in data/ds-jobs-*.jsonl — git-tracked, diffable, and the reason
a database outage never costs a day of collection. server/ingest-ds-jobs.js
then UPSERTs into Postgres: ds_jobs, ds_job_documents
(full-text indexed), ds_job_daily and ds_market_series. All of
it first-party — nothing in the data path leaves our own infrastructure.
Coverage caveat — the regional mix on this page reflects where our sources publish, not where the world's data jobs are. Remote-first boards skew heavily to North America and Western Europe. Read the regional chart as a statement about coverage, and read the ASEAN and Cambodia figures as a floor, never a total.