Tired of ads? Enjoy an ad-free experience by signing up.
  • Insights
    This article was written by a TIA community member. Insights pieces undergo the same rigorous editorial process that newsroom-produced articles have.
Ismail Afiff · · 7 min read

How we built an accurate autocomplete search feature at Traveloka

traveloka-autocomplete

Traveloka autocomplete feature

This article is part of Tech in Asia’s partnership with Traveloka, where we publish articles that feature the company’s valuable insights. Read more from Traveloka here.

Have you searched for hotels on Traveloka? I bet you’ve tried the autocomplete search feature, which gives users the most relevant hotels, regions, and landmarks as you type.

However, there were several challenges that we needed to address:

  1. The autocomplete feature should give the most relevant results using the least number of keystrokes.
  2. Since Traveloka deals with international users, the autocomplete feature must handle various languages and specific quirks. For example, in Thai, words are not separated by spaces. The island of Ko Samui is written as เกาะสมุย.
  3. The autocomplete feature should produce results swiftly. Otherwise, the user experience would not be enjoyable.

Here’s how we tackled these challenges.

1. Relevance score formulation

Users want to see only the most relevant results. Hence, the relevance score algorithm is key.

1.1 Classical way of text relevance scoring

The classical way of calculating relevance score is based on term frequency (tf), inverse document frequency (idf), and field-length norm (norm).

Term frequency means the more frequent a term appears, the more significant it is. For example, the text “Bora-bora Island” is more relevant than “Boracay Island” for query “bora.”

bora-bora

Bora-bora vs Boracay: battle of term frequency / Photo credit: corsarius_phil

Inverse document frequency implies that a word’s relative weight is related to the inverse of its occurrences in all documents, making more common words less significant than uncommon ones. For example, in the text “Eiffel Tower,” the word “Eiffel” carries more weight than the word “tower” because it is less common than the latter.

Field-length norm describes that the longer the text, the less significant it is compared to shorter ones. For example, for query “Jakarta,” the text “Jakarta” is more significant than “West Jakarta” because it is shorter.

2. Dealing with human languages

3. Creating a high-performance autocomplete system

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

Community Writer

Ismail Afiff

Ismail is a backend software engineer at Traveloka. Currently, he is developing search engine for Traveloka, helping millions of users to find their next adventures. Besides search-engine, he is also interested in Computational Science and Engineering.