Search by natural language

A product discovery piece exploring how AI could transform job search — and the approach that made design and data science genuinely useful to each other

Product discovery
10%
Rich search context >45 characters
47%
Medium search context 15 – 45 characters
43%
Low search context <15 characters

The background

Stepstone’s job search was built around a highly structured two-field form — one field for a job title, drawn from a normalised list, and one for location. Tightly controlled inputs, tightly controlled outputs. It worked within its own logic, but it meant job seekers could only express intent in a very narrow way: pick a title, pick a place.

That constraint had a real cost. Users couldn’t express what they actually wanted — the industry they cared about, the trade-offs they’d accept, the kind of environment they were looking for. The search bar as a surface was doing almost none of the work it could.

The opportunity

The goal was to enable natural language comprehension while giving priority to user intent — to redefine explicit intent-based job search by leveraging AI to surface highly personalised opportunities, minimising the emotional cost of searching and inspiring users throughout the journey.

But the challenge wasn’t purely technical. Even if the model could understand richer input, would users actually provide it? Stepstone’s users had been trained by years of keyword search to self-censor. The design problem was a behaviour change problem: how do you invite people to say more, without making it feel like effort, and without leading them towards answers that suit the algorithm rather than themselves?

Dual-tracking design and data science

Rather than handing a brief to data science and waiting for something to design around, we ran the two tracks in parallel from the start — each informing the other through a regular rhythm of shared sessions.

Data science explored what was technically possible: what entities could be reliably extracted, what confidence thresholds the model needed, what it could credibly deliver in results. Design explored what would genuinely meet user needs: which conversational patterns felt natural rather than leading, what interaction approaches would encourage richer input, and what the model needed to visibly do with that input for users to trust it.

Each track fed back into the other. Data science findings shaped what design could honestly promise users. Design research revealed what signals the model needed to make the experience valuable — and what it couldn’t afford to get wrong.

What we explored

The central design challenge was getting users to break a deeply ingrained habit. Years of keyword search had trained them to compress their intent into the shortest possible input. We explored two distinct approaches to encouraging richer, more natural expression.

01
In-context guidance with type-animation
Animated placeholder text that writes letter-by-letter inside the search field, modelling the kind of input we wanted users to produce. A second layer cycled through shorter examples to provide additional variety. The goal was to show, not tell — demonstrating the behaviour at the exact location and moment we needed it.
I’m looking for a senior product role where I can balance leadership with hands-on work
02
Guidance by topics
Interactive topic chips (Work/Life Balance, Salary, Career Growth, Location) that broke the task into smaller, focused pieces. Each chip opened a dedicated input with its own type-animation guidance. Rather than facing a blank field, users could approach their search sequentially — one dimension at a time.
Tell us about your work/life balance priorities...
Work/life balance Salary Career growth Location

What we learned

Type-animations were more powerful than expected. One user didn’t begin typing until the animated examples appeared — at which point they immediately wrote in full sentences with rich context, directly mirroring what they’d seen. The guidance worked precisely because it was in-place and in-form: the same medium the user was about to use.

Topic chips effectively solved the writer’s block problem. Users understood the chips immediately, and the focused input made the task feel approachable. The in-context animation within each topic made it feel nearly conversational — as if the product was asking a question rather than waiting for a command.

A nuance we hadn’t anticipated: simply selecting a topic, even without entering text, was itself a signal. Users expected that act alone to have an effect on results — a natural conversational pattern of stating a topic before being ready to elaborate. The model needed to treat it as intent, not a null input.

The critical failure mode was result quality. User buy-in at the input stage was consistently high. But when results didn’t visibly reflect the effort and context they’d provided, that buy-in evaporated quickly. The emotional contract with the user depended entirely on the model honouring what they’d said — which meant the design and technology had to mature together.

Where it led

Early live-test data showed meaningful movement: 10% of queries were high-context, with users providing approximately 20 additional characters of context beyond standard what/where searches. A promising signal that the interaction patterns were working.

Across the discovery period, the dual-track approach produced something more valuable than a set of designs or a trained model: a shared language between Design and Data Science for talking honestly about what the technology could deliver, what users would accept, and where the two needed to meet. That common ground shaped the four-milestone roadmap that followed and gave the team a clear basis for prioritisation.