SkillBridge: Canadian Labour-Market Skill Intelligence
Career advice is mostly vibes. "Learn Python" isn't a plan. SkillBridge answers what should I learn next from evidence: 900 real occupation profiles, four months of Canadian job postings, and a recommender that can explain its reasoning.
Canada publishes an enormous amount of labour data (occupational competency profiles, job postings, employment projections) and almost none of it reaches the person deciding what to study next. I wanted a system that takes "I do X, I want to do Y" and returns a ranked, justified list of skills to close the gap.
Merged five sources: ESDC's OaSIS competency profiles (900 occupations × 181 descriptors), Job Bank open postings from November 2025 through February 2026, LinkedIn Canada postings and extracted skills, COPS employment projections, and the NOC occupation taxonomy. The ETL alone removed 51,103 duplicate rows.
The proposal framed this as link prediction. But the OaSIS matrix is essentially complete, 162,899 of 162,900 cells observed, so "does this edge exist?" is always yes, and a model that answers yes to everything scores perfectly. I reformulated it as matrix completion: hide 10% of cells, predict the ordinal 0–5 rating, and ask the meaningful binary question instead ("is this descriptor core?", 14.1% positive rate).
Four components, evaluated as a ladder against real baselines: an occupation recommender (popularity → Jaccard CF → MF → BPR → node2vec); a shortage/balance/surplus classifier; a salary and regional-demand regressor; and an open-vocabulary NLP extractor that bridges free-text posting language to the OaSIS taxonomy. On top of those, a skill-gap recommender with cold-start fold-in, so it works for someone the system has never seen, plus a small web app over the exported model.
BPR won the ladder: nDCG@5 of 0.617, MRR of 0.821, ROC-AUC of 0.99 on the core-skill task. The salary model landed at MAE $12,088 and R² 0.64 against a target of R² 0.4. The shortage classifier hit weighted F1 0.76. The hybrid NLP extractor reached micro-F1 0.37 against a hand-built gold standard.
The surplus class in the shortage classifier is bad: F1 of 0.04, because there were only 17 test examples of it. The NLP extractor is markedly worse in French (F1 0.21) than English (0.39). Both are in the report at full volume, along with a fairness audit across TEER skill levels, because a model you can't state the failure modes of isn't finished.