How Airbnb uses proximity signals to personalize without relying on individual user history.
By: Wei Jiang, Bin Xu, Bharathi Thangamani, Weiwei Guo, Sundar Srinivasavaradhan, Tracy Yu, Huiji Gao, Swapnil Ghike, Michael Kinoti
Great personalization starts with knowing your user. But what happens when the user is a stranger?
A significant share of Airbnb users arrive without a login, without a recent search history, or without any prior booking — especially those landing from paid advertising or organic search. For these users, the ML models that power search ranking, destination recommendations, and marketing landing pages face the cold-start problem: there is little to no user signal available.
In this blog post, we describe Proximity Features — a new class of ML features that addresses both the cold-start problem and evolving privacy constraints simultaneously. The core hypothesis: aggregate activity patterns within a proximity bucket can provide useful signals for cold-start personalization without relying on individual user identity. By grouping nearby users into buckets and aggregating their collective signals, we can construct rich, location-aware features available immediately for any user whose IP can be geo-located — no login, no prior history, and no persistent user identifier required.
The challenge: serving users without individual signals
Personalization at Airbnb has historically relied heavily on a user’s identifier to link together past searches, bookings, and browsing patterns. This works well for returning, logged-in users. But on marketing and acquisition surfaces, many users are logged out or visiting for the first time, and standard user identifiers are unavailable. In other cases, users are logged in, and standard user identifiers are available, but the features derived from them are stale or too sparse to carry meaningful signals.
Privacy regulations compound the challenge. Under GDPR, non-essential tracking for marketing personalization requires an appropriate legal basis and, in many contexts, user consent. Separately, browser restrictions and the deprecation of third-party cookies reduce the availability of cross-site identifiers. Any feature system targeting this traffic must work without relying on persistent user identifiers for model inference.
The result: users on these surfaces were receiving little to no personalization.
The solution: proximity features
The proximity key
The key insight behind Proximity Features is geographic correlation. Users browsing Airbnb from the same city or neighborhood tend to search for similar destinations, prefer similar price ranges, and often share travel patterns and local context. A user in Seoul is far more likely to book Jeju Island, the largest island that’s part of South Korea, or Osaka, Japan, than is a user in New York. This pattern holds however much Airbnb does or doesn’t know about the individual user.

We operationalize this insight through a proximity key: a compact group key representing a local geographic cluster of approximately 1,000 users. The key encodes a quantized latitude/longitude tile and, for densely populated areas, an IP hash bucket index within that tile — fine-grained for a city center, coarser for a rural location. In either case, the key identifies a neighborhood, not an individual.
This proximity key functions as an aggregation key analogous to user_id: any ML model that currently conditions on a user identifier can instead condition on a proximity key to serve cold-start users (e.g., logged-out, anonymous, or first-time). Even logged-in users with meaningful historical information available can benefit from the additional information that use of the proximity key enables. However, unlike logged-in user information, the features associated with a key reflect the collective behavior of the local group, not the actions of any single user. No persistent user identifier is required at inference time.
So what are the actual features? Once users are grouped into a bucket, features are computed daily across three categories: short-term engagement (e.g., top destinations, room types, and median prices from recent searches); long-term booking patterns (e.g., booked destinations and travel party signals); and aggregate bucket metadata (e.g., bucket size and geographic density). A user who has never searched on Airbnb gets a feature vector derived from approximately 1,000 nearby users, effectively borrowing signals from the crowd.
Adaptive clustering
Computing proximity keys requires grouping all global users into stable buckets of approximately 1,000 each. The challenge is that the geographic density of Airbnb users varies by orders of magnitude: a single coordinate near a major airport can represent vastly more daily users than a rural location with only a handful. A fixed geographic grid (like a standard geohash at one resolution) fails at both extremes — too coarse for cities, too fine for rural areas.

We solve this with a two-phase adaptive clustering algorithm. For any coordinate dense enough to fill a bucket on its own, we subdivide further using IP hash buckets, which preserves fine granularity in city centers. For the remaining coordinates, we apply multi-pass coarsening, progressively widening the geographic tile until enough users accumulate. The result is a zoom-adaptive map: fine-grained tiles over major cities, coarser groupings in rural areas. This runs efficiently at global scale on geo-IP coordinate data.
Proximity keys are also stable over time: the key partition bootstrapped in 2023 has remained valid in production without re-clustering, with a daily refresh handling new IPs and coordinate shifts. At serving time, a user’s IP is resolved to a proximity key in real time and features are fetched from a distributed key-value store. The lookup is a soft dependency. If it times out, the model proceeds without it, so personalization never blocks the core request path.
Privacy design
Privacy is built in at every layer. Every feature reflects a group of ~1,000 users, not an individual. The clustering input excludes users who have not provided consent. Geo-IP coordinates are coarse and group-level, aggregated around population centers rather than specific addresses. The pipeline is also integrated with consent management and data-governance deletion controls throughout.
The benefits: what proximity features unlock
Proximity Features have now launched across multiple Airbnb surfaces, with measurable results from production A/B experiments.

Marketing landing pages. Pages where users arrive from paid advertising and organic search have especially high cold-start rates. Before this work, the listing recommender had severely sparse feature vectors for much of this traffic and served static listing cards. Adding proximity features gave the model access to booking and browsing patterns from nearby users, enabling meaningful personalization for cold-start traffic.
Homepage AutoSuggest. For users without history, the AutoSuggest feature on the Homepage previously fell back to a static global list (e.g., Paris, Barcelona, London, Rome), regardless of where the user was browsing from. Proximity features replace this generic fallback with location-aware signals: a new user browsing from Beijing now sees Hong Kong and Tokyo instead.

The experiment showed directional gains concentrated among never-booked and dormant users on Airbnb, exactly the population proximity features are designed to serve. The diversity shift corroborates this: destinations like Kuala Lumpur, Dubai, and Jeju Island, which in the past had rarely been displayed to users in relevant regions, appeared consistently in the treatment arm.
Engagement emails. The same serving layer is being extended to personalized email campaigns using recent coarse location signals. Experiment validation is underway.
Conclusion
Proximity Features show that the cold-start problem does not require a persistent user identity to solve. By grouping users geographically and aggregating their collective signals, we can construct features immediately available for any user whose IP can be geo-located, without relying on persistent user identifiers. The design is intentionally reusable; any model conditioned on a user identifier can be extended to cold-start users by conditioning on a proximity key instead.
For the full technical details, see our paper accepted to the TSMO Workshop at KDD 2026. The paper includes the system architecture, clustering algorithm, privacy design, and complete experiment results.
If this type of work interests you, check out some of our related positions!
Acknowledgments
The authors thank the Relevance & Personalization, Marketing Technology, and Infrastructure teams at Airbnb for their collaboration, feedback, and support in bringing Proximity Features to production.
We also want to thank Hui Gao for his support in authoring this post during his time at Airbnb.
All product names, logos, and brands are property of their respective owners. All company, product, and service names used in this website are for identification purposes only. Use of these names, logos, and brands does not imply endorsement.