← All posts
Jev vs. Regex: The Title Is Not the Job

30 September 2026 · Francesco Mucio

Jev vs. Regex: The Title Is Not the Job

We classified 3,000+ Berlin data and AI jobs with Jev, an evaluation model, after a regex baseline. It won the hard cases. And the easy one too, despite our mistakes

databerlin.net files every Berlin data and AI posting into one category, so you can browse “Data Engineer roles” or “genAI roles” instead of one undifferentiated list. For a long time that filing was a regular expression. Then we decided to test Jev, an evaluation model that had just come out and classifies far more accurately. This is the anatomy of that move, and of the review that told us we had it backwards.


In the beginning we filed jobs with a regular expression

A regex matches a list of keywords against the job title. It also cannot tell that a “Process & BI Manager” is a data job, and it cannot see that a role titled “Senior Data Scientist” is, in its body, an LLM engineering job. But it is fast, it costs nothing, and doesn’t require an LLM.

We knew our classification problem, and a set of regexes was just good enough.

Then TypeSafe AI announced Jev. It does not work like a chatbot: you hand it text and a set of typed questions, and it returns a label and a number, not a paragraph of prose. Cheap and fast at scale was exactly the shape we wanted. We like trying new things, so we tried this one.

The title is not the job

The title often does not describe the job.

Take a “Data Operations Analyst” posting filed as Data Engineer, because that is the work. A reader browsing “Data Analyst” will not find it where they expect. Or a “Staff Software Engineer, Data Platform” posting that is a** Data Engineer** role, the title says software, the work says data platform. Or a “Data Scientist” posting filed as** genAI/LLM/Agentic AI**, because the posting is about building with models, not analysing statistics.** A company names the role one thing; the work is another**. That is why we file a job where the work is, not where the* title* points.

Two steps is better than one

At first we tried to do it in one pass: send every posting to the model and have it pick one of the site’s categories — or a “non-data” bucket that meant “drop this one”. This is a recommended practice: letting the AI tell you “I don’t know.” A safety exit in case the model cannot decide.

That gave poor results. Whether a job belongs on the board at all is a different question from which of nine categories fits it, and asking for both at once muddied the second answer.

It was also wasteful: we download all jobs from a board, then filter for the data and AI ones. A single classification step meant a model call for each role, including the ones we were about to throw away.

So we split the work into two parts.

A cheap regular expression now decides data-ai/non-data-ai by title at ingestion and drops the rest before they ever reach the database. The gate is allowed to be dumb. It matches keywords against the title and nothing else. And if a company couldn’t spare the ten seconds to put the word “data” in the name of its own data role, the gate will not guess for them. Writing the title is the cheapest part of hiring, and some of them still ship a “Growth Hacker” and hope for the best.

Only the survivors reach Jev, where the actual reading happens.

Keep it simple, Sonnet — or was it Opus?

Initially the classification was not perfect. A pile of postings landed in the catch-all “Other”, cases where Jev genuinely could not tell what they were.

But why? We were missing some keywords? Our category descriptions weren’t clear enough?

The reality was that for some roles the classification was easy. For others there were multiple possible categories. The pipeline had been built using Claude Code and it had decided to protect us: a confidence score on every label, a fallback to the old regex when it ran low, and a human review on top.

This is usually a good idea when you are dealing with a company budget or life changing decisions, but if you are just deciding if an Analytics Specialist role is more an analytics engineer or an actual data analyst you can probably survive with a misclassification. No rocket surgery here!

The model’s rule is max-wins: whatever its top label is, take it. One guard remains: if the top label is the catch-all “Other”, it must beat the runner-up by at least 0.10, otherwise we take the runner-up.

No review queue. No keyword override. Everything we had added on top of the model got deleted, and the pipeline got better each time we deleted something.

The experiment that told us we were fooling ourselves

Discovering that we were overcomplicating things wasn’t easy.

What put us on the right track was a nagging question: was Jev actually better than our rules, or were we just victims of the hype around this model?

So we put them head to head. We ran both classifiers over every active posting and kept only the disagreements — the cases where a rule fired and landed on a different category than the model. That gave us 173 disagreements out of roughly 1,000 postings.** Jev won the hard cases**. The rules won the trivial ones.

Three independent LLM reviewers judged all 173, one case at a time, model against rule or a tie. The results were peculiar. Jev was right in two of three. The regex, in one.

JEV (the model)   109  63%
Rules (ours)       60  35%
TIE                 4   2%

And where each side won mattered more than the split. The rules won almost exclusively on literal, canonical titles: “Data Scientist”, “Data Analyst”, “Data Engineer”, “Head of”, “Director”. A substring match simply cannot miss those. The model won the rest.

That was the interesting part, so we went and looked. And the flaw turned out to be ours, not Jev’s.

We had handed our reviewers, lazy sons of an AI, a fragment of each posting: initially only the title, then the title plus the first few hundred characters of the description. The median description on the board runs about 5,400 characters, and it opens with the company blurb, which tells you nothing about the role. Jev was reading the whole posting. Our reviewers had read the part that does not matter and called the regex better.

We re-judged those cases against the complete description. And they flipped. Now the verdict across all 173, once everyone was finally reading the same text, was that Jev was always right.

Where the regex lost, in a few examples:

  • on seniority and leadership, flattening “Senior Director ML Engineering” into an individual-contributor AI/ML seat because they saw the “ML” first and ignored the “Director”;
  • on non-data roles carrying data or AI tokens: “Database Support Engineer”, or “QA Engineer – Core Database”.
  • on roles whose title had the “right words” despite being something else: a “Senior Project Manager, Data Platforms” filed as a Data Engineer, because the regex read “data platform” and never read “Project Manager”.

This also told us that maybe, maybe, maybe we could just trust Jev.

The taxonomy grew, only once for now

Jev also allowed us to split what used to be a single “AI/ML” category into two:** genAI/LLM/Agentic AI** and classic** AI/ML**. The new one needed a definition, not a keyword list.

Counting genAI and LLM words in job titles surfaces about 5% of postings. Counting them in title plus description hits 80–90%. In 2026 data and AI postings those words are everywhere, so they separate nothing. The category only earns its place when it means* building or shipping* the LLM, agent or RAG system, not* mentioning* it. Could we have done this with regex only? Probably not.

Today genAI/LLM/Agentic AI is the second-largest bucket on the live board, ahead of plain AI/ML: 235 live postings against 83 AI/ML and 52 Data Scientist.

We also went back through the big “Other” category to see if there are other role clusters, but we weren’t lucky. It seems that today every role can be a data or AI role.

Closing words

The board carries** 1,018 live postings** while we are writing.

Live (1,018)
Other                 323
genAI/LLM/Agentic AI  235
Data Engineer         130
AI/ML                  83
Data Analyst           71
Leadership             64
Data Scientist         52
Product Manager        43
Analytics Engineer     17

Behind it sits the archive:** 3,072 classified postings** — everything the board has seen since we started three to four months ago, live and closed alike.

A handful of just-ingested postings sit outside that count, waiting for their first pass.

Seniority is still assigned by keywords, and that is fine. Seniority lives in the title and in the role metadata (if they exist), where a regex is a perfectly good reader. Category does not.

An evaluation model like Jev beats a keyword list on this task. And it lets us reclassify the whole board tomorrow.

Its advantage lives entirely in the text it is given, and it turns against you the moment you judge it from a fragment that stops before the part that matters. But you probably already know this. The question is, does your coding agent know this too?


This post first appeared in the select* from newsletter.