← All posts
Teaching a Job Board to Find Skills It Didn't Know to Look For

8 September 2026 · Francesco Mucio

Teaching a Job Board to Find Skills It Didn't Know to Look For

We tried GLiNER2 zero-shot extraction to discover new skills in 1,769 Berlin job postings. It found real signal, and about ten times as much noise.

databerlin.net tags every job with the skills it asks for: Python, dbt, Kubernetes, that sort of thing. The dictionary behind that is hand-maintained, and hand-maintained things fall behind. Here’s the experiment we ran to fix that, and what it actually took to make it useful.


Start with the problem. Why did the skill dictionary need help?

We already had a miner, and the algorithm behind it is simple: regex-tokenize every job’s title and description, count how many distinct jobs each token shows up in, keep the ones that clear a frequency threshold and pass a shape test: has a digit (k8s, python3), starts with +/# (C++), all-caps 2–8 letters (ETL, AWS), or CamelCase (BigQuery). No model, no embeddings, no notion of meaning. Regex and counting.

That was enough for a long time, because most things people list as a “skill” do happen to get spelled in a technical-looking way. Tool and framework names are capitalized or acronym-ized by convention, and the heuristic just rides that convention.

Where it breaks: anything without a shape tell. “Vector databases,” “prompt engineering,” “reasoning tokens”: multi-word, lowercase, spelled exactly like ordinary English. The frequency counting still works on those; they show up often enough. It’s the shape filter that throws them out before they’re even considered, because nothing about how they’re spelled marks them as technical.

That’s the gap, and it’s growing. What people list as skills in 2026 postings increasingly looks like that, not single capitalized tokens.


So you reached for a different kind of model, one built specifically to pull named things out of text (that’s what “NER,” named entity recognition, means). What is GLiNER, and how is that different from just asking an LLM to extract skills?

GLiNER doesn’t write anything. You give it a label, “tool or technology,” plus a block of text, and it marks which words in that text match the label. A highlighter, not a writer. An LLM does this differently: you’d prompt Claude to read the text and write out a list, then hope what comes back is complete and doesn’t paraphrase or invent anything. GLiNER skips the writing step, which is also why it’s smaller and cheaper: small enough to run on a laptop CPU, fast, and it gives the same answer every time for the same input. No generation means no randomness.

The trade-off is what you’d expect from a small, narrow model: no world knowledge, no reasoning about a case it’s unsure of. It can’t tell “Tiger Global” is an investor, not a technology. It only knows the phrase sits where technology names usually sit in a sentence. An LLM would probably get that right, at maybe 50x the cost.

This whole approach is called zero-shot: you don’t train it on examples of what a “skill” looks like beforehand, you just describe the label in a sentence or two and it goes looking. That’s the appeal: no dataset to build. It’s also the main limitation, as we’ll see.

We used GLiNER2 from Fastino Labs, specifically their 2.5 release. Older versions of GLiNER found matches by generating every possible highlight of the text (every start point, every possible length) and scoring each one against your label. That’s expensive, and it caps how long a match can be. Version 2.5 changed the approach: instead of enumerating candidates, it just predicts, per word, “is this where a match starts” and “is this where it ends.” No list of candidates to score, so it stays cheap even on long text. We used the smallest size, gliner2.5-small-v1, sized for CPU.

First test was 15 jobs. Then 200. Then, once I trusted the filtering, all 1,769.


Could you fine-tune it to get better results?

Not quite. We never retrained the model or showed it labeled examples; nothing was “learned.” What we tuned was the schema, the label description handed to the model each time we run it.

First pass, the label was just "skill", wide open, pulled in benefits copy and job titles along with real tools. Second pass, I narrowed it to "tool_or_tech" with explicit negative instructions baked into the description: “not a company name, not a job title, not a benefit.” That’s schema engineering, the GLiNER equivalent of prompt engineering: reshaping what you ask for, not retraining what the model knows.

It helped (visibly less junk in the results) but it never got the output clean. A company name still looks like a proper noun sitting next to technology terms, no matter how you phrase the label. That’s the ceiling of zero-shot: you can aim it more precisely with words, but you can’t teach it actual judgment without real training data, which needs way more setup than a maintenance script deserves.


What did the raw output look like?

Noisy. Real stuff like TabPFN, LangSmith, Data Vault 2.0, SGLang mixed in with company names, investor names, and German HR boilerplate, all extracted with the same confidence.

The model doesn’t know the difference between “we use Snowflake” and “backed by Tiger Global.” Both are proper nouns near technical-sounding text. It also happily extracted benefits copy (“height-adjustable desks,” “meditation budget”) and gender tags like m/w/d, because those show up in the same sentence position as real qualifications.


How did you deal with that?

Layered filtering, mostly reactive. Run a batch, see what leaked through, add a rule. Company names cross-checked against our own companies.yml. Job-title suffixes. Degree requirements. Benefit language. Gender tags. Department names. Nothing designed upfront. It grew every time something dumb got through.

One bug actually mattered: checking whether a candidate was already covered by an existing skill’s regex used re.search, “does this pattern appear somewhere in the string.” Wrong direction: “Claude Code” got silently swallowed because our “Anthropic Claude” pattern matches the substring “Claude.” Switching to re.fullmatch (the whole candidate must match) fixed it. A small bug, but it would have killed exactly the new candidates I wanted.


Give me the numbers. How much of what it found was real?

Across the full run: about 1,600 candidate strings worth reviewing. After triage so far, batch by batch, ranked by how many distinct jobs mentioned each one, we’ve made 1,043 keep/reject/merge decisions. 944 rejected. 92 kept. 7 merged as aliases into existing entries. That’s roughly a 9% hit rate on raw model output.

Most of the rejects fall into a few buckets: company and investor names, generic business metrics that show up in job descriptions but aren’t skills (ARPU, CAC, NRR), job-title fragments, and near-duplicates the fullmatch check still missed because the phrasing differed just enough: GCP GBQ vs BigQuery, ML/AI restating two entries we already had.


Any candidates that looked right but weren’t?

Yes. The one that stuck with me was “automated decision systems.” Eight jobs, looked like a real compliance/ML-governance term. Then I checked: all eight came from the same company, same paragraph. One employer’s boilerplate, counted eight times because we dedupe per-job, not per-employer. Job-count alone isn’t signal; you have to actually look at which jobs, not just how many.


Was the review itself automated?

No, and it shouldn’t be. The model is a candidate generator, not a judge. Every batch got reviewed by hand, 20–50 at a time, sorted by job frequency. My own noise filters caught some real skills early on too (JSON, VS Code) and needed a manual audit pass to restore them. The model proposes. A person decides.


Is this running automatically now?

Deliberately not. It’s not wired into run_all.py. This stays a manual, periodic tool: run it, review a batch, move on. A decision log (skill_candidate_decisions.yml) means anything already ruled on doesn’t resurface next run. Without it, we’d be re-litigating “Kununu is a company, not a skill” forever.


Worth doing again?

Yes, but the lesson isn’t “NER finds skills.” It’s “NER finds noun phrases near other noun phrases, and about one in ten of those happens to be a skill.” That ratio is fine if you have a cheap way to review in bulk and a memory that doesn’t forget what you already rejected. Without both, you just get a bigger pile of reviewed: false.


The skills you see on databerlin.net are still a mix of hand-curated and model-suggested; every single one reviewed by a person before it ships. If you spot one that’s wrong or missing, that feedback loop is exactly how the dictionary gets better.