DataCareer App
Classified 7,039 Job Listings by Skill Evidence, Surfacing 1,134 Hidden Data Roles
0
Listings Classified
0
Hidden Data Jobs Surfaced
0%
of Listings Were Hidden
Executive Summary
DataCareer App sourced Australian data-role listings by matching job titles against a keyword list, a method that missed roles like Business Analyst or Reporting Officer that are data jobs in substance, while letting administrative roles with "data" in the title through. As Data Quality & Insights Analyst, I was asked to build a defensible basis for how listings were actually classified. I reframed the question from what a job is called to what skills it requires, building a weighted skill-scoring framework from 34 skill keywords drawn from listing descriptions, each scored by how strongly it signals genuine data work. Applied across an extract of 7,039 listings, the framework surfaced 1,134 hidden data jobs (16.1%) that title-based sourcing had missed entirely, while flagging 1,076 listings (15.3%) as noise. Roughly one in five real data roles had been invisible to the previous method, and the classification became the basis for how the platform now categorises its job database.
Context
DataCareer App is a data-driven platform serving client organisations with labour market and career insight. The platform sourced Australian data-role listings from job boards including Seek and LinkedIn, refreshing daily, ingesting any listing whose title contained a recognised data keyword. I joined as Data Quality & Insights Analyst for a defined engagement from October to December 2025.
Challenge
Titles are an unreliable guide to job content. Roles advertised as Business Analyst, Reporting Officer or Insights Specialist are frequently data roles in substance and were being missed entirely, while some listings containing "data" in the title were administrative roles with no analytical content. The database was therefore both incomplete and imprecise, and the market statistics the platform published understated the real size of the Australian data job market.
Strategic Approach
Phase 1
Reframing the Question
Rather than expanding the keyword list, which would have added noise without addressing the cause, I reframed classification around the skills a listing's description and requirements actually demonstrate, rather than its title.
Phase 2
Building the Scoring Model
I built a weighted skill-scoring framework, a dictionary of 34 skill keywords drawn from listing descriptions, each assigned a point value from one to five by how strongly it signals genuine data work: SQL and A/B testing scored highest, generic tools like Jira and Confluence scored lowest. I chose a transparent rule-based model over a trained classifier because no labelled dataset existed to validate against, and the product team needed logic they could inspect, defend and adjust themselves.
Phase 3
Classification and Validation
I normalised the description text so keyword variants resolved to single terms, counted the distinct skills each of 7,039 listings matched, totalled the weighted score, and cross-tabulated that skill evidence against whether the job title was an obvious data title, resolving every listing into one of three classes.
Quantifiable Outcomes
- 4,829 listings (68.6%) confirmed as visible data jobs.
- 1,134 listings (16.1%) identified as hidden data jobs: genuine data roles that title-based sourcing would never have surfaced.
- 1,076 listings (15.3%) reclassified as not data jobs at all, despite title-based matches.
- Roughly one in five real data roles had been invisible to the platform's previous sourcing method.
Qualitative Achievements
- Extended the analysis into breakdowns by industry and state to show where hidden roles concentrate.
- The classification became the basis for how the platform categorises its job database, including a user-facing filter for hidden data jobs.
- Delivered a rule-based model the product team could inspect, defend and adjust themselves, rather than an opaque trained classifier.
Expanding a keyword list treats the symptom. The database wasn't missing keywords, it was asking the wrong question: what a job is called, instead of what it actually requires.
This engagement is the clearest evidence of how I approach a messy classification problem: trace where the method is actually breaking down, build something transparent enough for a non-technical team to trust and adjust, then validate it at scale. It is the analytical counterpart to the automation work at DashboardWorx.