#34: From Consumer Memory to Semantic IDs: Generative RecSys for Quick Commerce with Raghav Saboo
Note: This transcript has been generated automatically using OpenAI's whisper and may contain inaccuracies or errors. We recommend listening to the audio for a better understanding of the content. Please feel free to reach out if you spot any corrections that need to be made. Thank you for your understanding.
It's not just getting more clicks or getting more engagement, it's actually understanding what's healthy for such a marketplace to grow in a way that benefits all of the three stakeholders in this flywheel.
Within e-commerce, items really carry a lot of meaning.
I don't think our current evolution of user experiences of how we show search results and how we show recommendations has evolved recently, like all apps look the same.
Our surfaces, our user experiences are non-malleable enough right now to truly be unified.
Semantic IDs, even though they're content-based today for our construct, they work well for personalization tasks like recommendations.
The other use case that we wanted to test, which was to see how well they work for our search, was using semantic IDs for query reformulation.
It's not that agenic systems themselves are a new system, but rather a user experience question.
Hello and welcome to this new episode of Recsperts, Recommender Systems Experts.
In today's episode, it will be everything about generative recommendations.
And this time again, we will be talking about recommendations, ranking, search, and personalization as the overarching broad scheme in a domain that we have been talking about in the past already.
And the not too distant past.
And yes, you are right.
This time we are talking again, but slightly differently as well, about the delivery domain and marketplaces.
For this episode, I have invited another colleague of mine, Raghva Saboo.
Hey Raghav, welcome to the show.
Hey Marcel.
Yeah, thanks for inviting me.
Very excited to be here.
Yeah, great that you joined.
This won't be the only time that the both of us will be somewhat on stage, but also towards the end of this month, the 28th of September, as of this year's RecSys 2026 taking place in Minnesota, we will also be on stage joined by another colleague, Pawel Kams from Volt.
The three of us are going to hold a tutorial on recommender systems in delivery platforms, challenges, solutions, and learnings.
So for everybody who is listening to this and maybe already on their way to RecSys or soon on their way to RecSys, you don't want to miss out and check that abstract.
And we will also be talking about this more broadly towards the end of today's episode.
And we hope we do our best to tease you and see a lot of familiar and also unfamiliar faces in our tutorial to learn more about recommender systems in delivery platforms.
But back to our guests, back to Raghav.
My guest today, Raghva Saboo is a staff machine learning engineer working for DoorDash.
And he brings more than five years of experience in this position.
So with this being said, I know that I said it already earlier, but Raghav happy dash-versary again.
Yeah, thank you.
It's been a while, but it's been fun.
That's great to hear.
And yeah, it's not only be fun, but also been very impactful work that Raghav has been performing.
And not without a reason, he has become the tech lead for search and personalization machine learning in the new verticals business line, which more concretely is the domain where we at DoorDash do grocery convenience and retail, because yes, we are not just a restaurant delivery platform.
We are much more than that.
He was previously with Amazon where his main focus was on developing the first set of large language models and distilled models for new language launches for the Alexa AI.
And prior to that, he was a machine learning consultant.
Raghav holds a masters from Duke University in statistics, machine learning and econometrics, and also a bachelor of engineering and master of engineering in surprise, surprise, chemical engineering.
So maybe one of the fun questions for today's episode might go there from the Imperial College in London.
Raghav has also been very active when it comes to publications, not just in DoorDash's blog, but also with lots of publications at, of course, Recommender Systems Conference, but also the Wisdom, SIGIR, CIKM and KDD.
But before we dive deeper into all of these fascinating, exciting and impactful topics, I definitely want to hand over to my guest for today.
Tell us more about yourself, Raghav.
How did you get there where you are today and what has made you shift your career from chemical engineering, building or designing plants for ibuprofen into becoming such an impactful and great colleague and tech lead for search and personalization ML at DoorDash?
Yeah, no, thank you, Marcel, for inviting me to this podcast.
I've been an avid listener of the podcast.
I'm very excited to be part of it now.
Definitely hoping that we can cover a lot of the topics that we want to share at RecSys this year as well as part of our tutorial.
And yeah, chat about a lot of work that we've been doing at DoorDash, Wolt and Deliveroo together as well.
So thanks again for inviting me.
What has brought you into Recommender Systems, into personalization, and how came the career change from chemical engineering into designing Alexa, AI and working on probably the first LLMs and then stepping into personalization and becoming very successful there over the course of the past five years?
Yeah, it's been an interesting journey in terms of moving from chemical engineering to the ML and AI space, but it hasn't been as different as a domain for me, interestingly.
So when I started working as a chemical engineer after graduating, I actually started working on building automation loops for these chemical plants across Europe.
These are fairly small chemical plants, but they could be run remotely.
That was at least the goal of these plants.
And automation within these plants has been decades old.
So creating feedback loops and building some sort of PID processes to control these plants was an area of chemical engineering that was studied fairly well.
But through that work, I actually started getting more interested in machine learning and statistics more generally, because I started to read more about it during that time as well.
During my chemical engineering training, there was very little that was covered.
And that's what made me ultimately make the move and decide to do a degree and move to the United States and start applying some of this work as well.
And over the past few years, it's been kind of like a journey towards figuring out which part of machine learning and AI I find the most fascinating.
Today, working on search and recommendations feels like the natural fit for me in many ways.
It ties in a lot of my interests, and I think it's a very rich space with a fascinating set of problems to solve.
Yeah, that's kind of been my journey.
Interestingly, after chemical engineering and doing stats, my initial goal was to do energy policy work with machine learning.
And I did do quite a bit of that, working with the World Bank to figure out how renewable energy adoption and electric grid upgrades in South Asia could be accelerated.
But just working within the policy space, I realized the speed of impact is much slower.
And the reason I moved into consulting was I didn't know which part of the industry I wanted to actually apply this work in.
I think this career path is probably going to help me navigate and find an area that I enjoy working in the most.
Understand.
So it's fun that as you say this, I start to see that we definitely have things in common in terms of our early career.
Because for me, the theme is a bit like we started in the physical world and we somewhat moved into the virtual world.
Not to say that a lot of what we do at Doordish, Volt and Deliveroo is actually having an effect and making a difference in the physical world.
But when we work on certain things and it's very virtual and sometimes highly abstract.
Because when I was starting my career and even when I was doing my bachelor's and master's, I was very much into production engineering and logistics.
And then throughout that, developing for example, the arrangement of different material placements and buckets in a shelf.
I was designing a linear programming setup for this and checking like, okay, how can you improve picking efficiency while maintaining certain ergonomic ways for logistics employees.
And then I was also actually working for an energy trading company as a working student.
But then also over time, I somewhat got interested in machine learning and what you can apply to.
And then also like got the chance to apply this to recommender systems, which I'm sticking to now for almost 10 years.
So when you were working on large language models at Amazon for the Alexa AI, have you already seen what is probably coming with it?
So 2022, chat GPT, nowadays agents, generative recommendations.
Was this providing you with a glimpse into the future of what might happen after voice recognition and voice understanding and what you were working on there?
So what have you been doing there and how much has it probably informed you of what is coming?
Yeah, I think the Alexa AI experience was definitely transformative in my thoughts of where a lot of the technology transformers had just been recently introduced and the application of it and language models was like the first wave.
And whilst I was at Amazon, bird style models had become quite commonplace.
And the first set of areas of focus for a lot of these companies was to figure out how do you build a large teacher model, multi-billion parameter encoder models and use that to bootstrap tiny birds, let's say.
Right.
Which is fascinating to me because these would be what are probably today considered as small language models.
But even then, we were training these large teacher models as large as like 9 billion parameters.
And the goal was that using these large teacher models for resource rich languages and using that to then train a smaller student model for resource constraint languages and effectively do transfer learning in the language domain.
And these smaller language models were maybe 17 million parameters or so, so like around 20 to as small as 5 million parameters.
All right.
Yeah, through that work, we definitely saw the value of pre-trained models and the advantage that transformer-based architecture brings to these tasks.
Before these models, a lot of these companies similar to Alexa would be using very traditional language modeling architectures.
And this was definitely a step function in how quickly we could start launching our new regions and new languages.
So yeah, back to your question, I definitely think we saw the early indications that this could effectively be a scaling law kind of question where if we add more data, increase the parameters, these models will generalize.
I think I was not prescient enough to understand was that this would even translate beyond language tasks and that language or tokens rather as a primitive construct for intelligence could scale so well.
Today we have general purpose models that are trained on multiple tasks that are beating benchmarks everywhere.
And that was not obvious to me at that point.
Let's step into the work that you are doing actually nowadays at DoorDash as a tech lead for search and personalization ML there.
And here also for our listeners, a short reference back in episode 32 where I talked with Sasha Fedintsev, we have talked more about the restaurant side of the business.
Today we will complement this with the new vertical side of the business.
And before saying that, maybe just also a little bit of background about DoorDash, Volt, Deliveroo.
So DoorDash is a local commerce platform and delivery marketplace which is operating in more than 40 countries with over 56 million monthly active users within the group.
In 2021, DoorDash has acquired Finland's Volt.
And four years after last year in 2025, DoorDash acquired Deliveroo, which was a UK-based delivery service.
And DoorDash is also part of the Nasdaq 100.
And nowadays we are growing together as one single company that operates across the world, across these brands.
And this is, I guess, also what we can say a lot of what is part of our daily work to build this future of one single platform.
And many people still think that what we do is deliver your pizza, deliver your sushi, and you order, let's say, something for your hangover or whatever.
Yes, we do that.
But as I already said- Midnight munchies or cookies.
Yeah, exactly.
What is it?
Mac and cheese?
Mac and cheese.
But like I was just saying, like midnight cookies are pretty common on DoorDash.
Okay.
Okay.
Yeah.
But we do much more than just delivering food, but we also deliver groceries.
You can get your new iPhone from DoorDash, but you can also buy alcohol or stuff and food for your pets and whatever you think about.
And I have the pleasure to sit in front of that person who is basically thinking in and out about how to personalize this domain for our users in all different facets.
Before we dive into that, Aragaf, could you tell us more about this operating model of a multi-sided marketplace and also how this new verticals domain positions itself within the setting?
Yeah, that's a great question.
I think like multi-sided marketplaces have become much more common, especially since COVID, right?
Like as a lot of the physical economy has moved online, COVID definitely accelerated a lot of that.
DoorDash is in many ways no different to a beneficiary of that as well.
And when we look at that bridge from the offline physical commerce to the online commerce, restaurants was the first area that a majority of these delivery companies started with.
DoorDash has grown quite significantly and beyond that now where we are really tackling the last mile problem, I would say, or anything that you could think of that you would want to potentially purchase within your kind of locality or beyond, right?
And new verticals, as we call it on DoorDash, started off with groceries, expanded into general retail, right?
And today we kind of have multiple, even small and medium businesses that are local mom and pop shops, maybe your local tea sellers, and all of them are on the platform.
And our goal is really to create a very rich marketplace where the benefit of our consumers, obviously, through the last mile delivery problem where we provide the conveniences of a delivery platform and a discovery platform in many ways.
And then for our merchants to reach a wider consumer base.
But there's also a third component of this flywheel, which is our dashers, in large, like our fulfillment network, right?
And so when you look at multi-sided marketplaces and how they interplay with discovery surfaces on an app like DoorDash, and these discovery surfaces are usually search based or recommendation based, there's a lot of considerations that you have to make for any product change, any model change that you're making.
And that's what makes it very fascinating and rich, right?
Because it's not just getting more clicks or getting more engagement.
It's actually understanding what's healthy for such a marketplace to grow in a way that benefits all of the three stakeholders in this flywheel, right?
If our consumers are able to find items that are more niche, more suited to their preferences, we need to be able to surface those items in a way that also has some fulfillment, quality, and assurance on getting that item, right?
That we don't want to frustrate our users and not getting the item that they expected.
At the same time, we need to also balance how our merchants are able to also demonstrate their core brand values and balance that with purely personalizing for the consumer.
So yeah, it's a pretty fascinating space.
And I don't think some of you have solved all of these problems, but the fact that these problems exist is what keeps me really excited.
Jens Stoltenberg That's great to hear.
And a lot of the problems that this area brings or challenges for personalizing, especially keeping the consumer in mind.
And I guess we could definitely say that even though this is a multi-sided marketplaces in the work that we both are doing on different ends of the platform, we are like traditionally like very consumer centric.
This is also kind of reflected, I may say, by the published work that you with your peers published over the last one to two years.
And there was something more concretely that I would like to start with, and then we can go into the more current topics and also dive into this very interesting aspect of unifying search and recommendation, something that is upcoming also at this year's RecSys conference with the first workshop on unified search and recommendation, where you will be giving a presentation.
So something I'm looking at right now is based on a talk that you and Sudip Das, the head of new verticals, ML and AI at Doordash gave at the KDD 2025 Paris workshop.
But let's turn it a bit differently because like we are the RecSys enthusiasts, so we will rather refer to the Gen AI workshop as part of last year's RecSys, where you were actually talking about how you leveraged LLM assisted personalization for solving three things that matter to our consumers, which are affordability, familiarity, and novelty.
Can you walk us through these three aspects and why and how they matter to consumers and how you leveraged LLMs in assisting the personalization here?
Yeah.
So I would say like the use of LLMs in that case, right, was not the core kind of premise of it, but it was the fact that, you know, these are key objectives at a high level for a healthy marketplace.
The use of LLMs has helped us really leapfrog a lot of the problems that would take maybe multiple quarters for us, not in terms of just development, but also in terms of like the ML architecture and feature engineering, you can call it more traditionally.
But yeah, maybe I can first describe like the three objectives at a high level where people who have not read the blog posts that we also did or did not attend the talk.
So when we talk about familiarity, affordability, and novelty, familiarity from our kind of recommendation system parlance is you want to surface items that you already love and trust, right?
That is what helps build confidence in a platform for a consumer.
You want to feel like the app is working for you and reducing the friction.
And on an app like DoorDash or any e-commerce platform app like ours, you want to be able to get to the items that you want to order quickly.
The other aspect of it is affordability.
So affordability is really about meeting our consumers where their preferences lie.
Right.
And this is not purely about pricing, but also about surfacing the right deals at the right time for consumers.
And it should feel like the marketplace is there for consumers as an option, and it shouldn't feel like an expensive option for them or an option that is less relevant from the affordability front in comparison to like our physical stores.
The third aspect is novelty.
And this is, you know, in our kind of RecSys speak, similar to exploration, right?
So familiarity brings you the items that we have very high confidence on that you would want to order at any given time, given, you know, the intensity you've recognized.
Novelty is about surfacing, again, relevant items as far as possible to you or learning about you over time, but doing so quickly and still in a way that does not add friction.
Right.
So for a marketplace, it could be introducing complementary items as you're building a basket, introducing you to new categories of items that we anticipate you may like based on other similar consumers.
The idea is now that you have like a framework to say, OK, you know, familiarity, novelty and affordability are, you know, three high level concepts that you want to balance the marketplace for.
What does it mean to like actually drive against these objectives?
Right.
And driving against these objectives has like in our framework and our stack multiple components.
It's not purely kind of optimizing on these as like a task for ML model or like a multi objective optimization through online experimentation.
It's actually building kind of like an ecosystem of components.
So what are the components?
I would say there may be five components at a high level.
The first set of components are really built around figuring out how do you create a content pool that is relevant for the pool of consumers that you have in the marketplace that could take the shape of, for example, collections on a store page or on any surface within an app collection of items that describe some sort of thematic intent.
So for a local marketplace app like DoorDash, it means detecting occasions that might be happening.
So you could have some collections that are themed around upcoming public holidays in the US, Thanksgiving, Valentine's Day, etc.
You would have also some pre-generated content that you would need to surface for specific shopping missions.
So let's say we know that you have pets and you need to train the pet to like training toys, training treats.
Similarly, within grocery, your weekly shopping missions, grocery produce pantry missions.
So you can distill a lot of this down into these intent-driven content.
And that's where we leverage LLMs quite extensively over the past two years is really scaling that out.
So previously, a lot of this was driven by our ops team, the operating team who work very closely with our merchants and with our business to figure out what is relevant for a platform like ours.
And with LLMs, we're able to pass a lot of that context over to agents that can generate the same content and actually do it at scale where it's personalized for each consumer as well.
So that's, I would say, like the first component.
And this first component is broken into two in some ways.
We have consumer-specific content.
And then we have merchant-specific content as well.
So you don't want a carpet seller to look like a grocery store.
But because we have so many niche merchants as well, we need to surface right categories, right intents to really provide value proposition of that merchant up front.
What does it mean like very concretely?
Like you were saying, there's this ecosystem of five components.
One of the components that you just are walking us through can be subdivided into this merchant perspective and consumer perspective.
So what is it in the end?
Like I feed an LLM with a lot of data about the consumer and then I ask for it, okay, what are relevant collections, whatever collection means in that regard for the consumer?
And then I store that with identifiers in a database that I could use for candidate generation or what is it used for and what are the artifacts that are generated as part of this?
Correct.
Yeah.
So at a high level, that is what happens, but it's a bit more nuanced than that.
And just talking about the inputs over here, right?
So as generative AI has picked up, the use of LLMs is pretty broad.
So LLMs being used as just single-term input-output generators to multi-term agents.
And the approach we've taken when building this ecosystem for such generative recommendation applications is to think about primitives that can be reused across these use cases, right?
So for something like generating content in the form of collections to represent consumers, we have built, you can call it a primitive around memory, right?
So memory is essentially, if you look at current literature, it's a weight for agents to obviously learn over time and do continual learning through taking notes in a very simplistic way, right?
So it's like markdown notes for your coding agents, but how does that translate to agents for apps like DoorDash?
It's like essentially letting LLMs consolidate its understanding around a consumer over time.
And we structure those through a concept of memory blocks, essentially structured domains or semantic domains around which we can define a consumer.
This could be something like, what are their dietary preferences that we've understood from their order history, whether they have pets or not, whether they have specific brands, they trust more than others, right?
And these are, again, different from regular features or profiles you can generate about consumers from a traditional perspective, you would maybe look at aggregations, embeddings, et cetera.
In the case of memory, it's actual natural language representations of consumers that again, LLMs for content generation can use in a single term or agents in the form of like Ask DoorDash, for example, product that we launched earlier this year, which is a grocery ordering agent and a restaurant ordering agent can also use to personalize its sections.
That forms kind of like a primitive as an input, right?
So now we have a way to represent consumers for all LLM use cases.
In this case, we're using that input for collection generation, which yes, then like, you know, based on our pipelines, we're able to for different intents, generate personalized set of collections that becomes part of like the retrieval pool itself.
Similarly on the merchant side, you know, we have the concept of merchant memory and we, you know, again, do a similar thing where we kind of build profiles for merchants based on different attributes or semantic domains that may help us categorize those merchants and highlight their unique value proposition.
So once we have that primitive, the content generation becomes essentially like LLM prompt engineering and inference problem.
But you can imagine, you know, that is just also a beginning, right?
Like you can start thinking of like, how do you represent that content?
So today, and this is also published on DoorDash's engineering blog posts, one of the ways that we represent these collections is through search terms.
Essentially, think of the LLM taking the understanding of a consumer, thinking of themes that might be relevant for the consumer and intents and shopping missions that the consumer may come with to the app.
When generating the content that's relevant for the consumer, it thinks of the searches that the consumer might do in order to get to the items that they would want to add to cart.
And then this is where we blend our personalization and search stacks.
So these search terms are resolved into items through, you know, embedding based retrieval on our search stack and then personalized through our ranking models.
There's already like one way that we're kind of seeing both of these stacks kind of emerging.
But yeah, so that's kind of like one way to look at content generation, right?
Now, there are natural evolutions of this, even if you look at current literature in the industry, right, representing items as semantic IDs, for example, and teaching small language models or language models about those semantic IDs, so that the language model is grounded in the catalog that is actually available at a given merchant or at DoorDash at large.
And then taking that further, where it's not just a single turn output, but a multi-turn and maybe a reinforcement learning loop where, you know, content is generated and we learn from the feedback on what works and what doesn't work for different types of consumers.
I think our listeners can get an understanding of this.
I, as an employee, had the chance to test it myself.
So I was creating this consumer profile, this consumer restaurants profile for my own user based on orders and associated metadata for those items and deliveries.
And I had to say, like, I was always pretty excited and also positively surprised by the nuances that this was able to grasp on weekdays and during lunchtime.
You are more likely to order hummus based dishes, bowls, salads, often with salmon and this and that.
It seems like you are a healthy eater, especially when you are at work.
And then on the evening, like early evening, it looks slightly different.
And the later the day gets, it might get even more unhealthy sometimes.
So there are definitely these once in a time, like burger deliveries around 10 PM or having a pizza.
But it was nice to see, like, how it grasps these nuances and maybe that already answers, but I want you, maybe, to add to this, why this is possibly superior to how it has been done before.
Because like my initial or maybe from a traditional point of RecSys architecture view, you might say like, okay, have your user tower, have your item, then you store whatever tower, whatever is your item in the specific problem setting and let them interact, let them create proper user and item condensed representations as embeddings and then do embedding based retrieval to solve for a relevance problem such as click probability, purchase probability, whatever, and then simply use those embeddings that you generate for users or for items.
What is the big difference of your LLM approach to that?
Is it really about having more nuances, having interpretability or what were really like these inspiring, exciting factors that you saw become very useful as a result of that work?
Yeah, I think that's a pretty pertinent question. I think generally, as a lot of industry practitioners are like looking to use language models and seeing where they perform well for things like recommendation systems, I think just from the concept of like traditional representation, like embedding based representations or both handcrafted features and also trained embeddings. Ultimately, I think the bottleneck of these representations come from two sources, in my opinion. One is the pure metadata that we have around a lot of the entities that we are trying to represent. So in the case of DoorDash, we have a pretty vast catalog and it's on the scale of like a few billion items if we look at items across all merchants that are not de-duplicated. Obviously, after consolidating those items, it may fall into the range of like a few hundred million items, but that's still sizable as a catalog to choose from.
And the other nuance that comes from that is the fact that a lot of the catalog that we surface are items that our merchants are providing data for. In that case, a lot of data can be incomplete. So whilst you would imagine that you would have a catalog with very clean, structured metadata about every product that we're selling on our platform, that is not achievable in many ways because of both the noisiness of the data that's coming in and also just the scale of the problem of cleaning that data. So what that leaves on the table with traditional models is the fact that whatever is not structured, we can't represent even through, let's say, if we use semantic embeddings and we just passed all that data as part of the semantic embedding input.
There are a lot of nuances that are not explicitly understood about that product, which is where I think LLMs provide an advantage where there's kind of like a reasoning component to it and representing that reasoning is, I think, the value that they bring.
So that's from the item side, from the consumer side as well. Ultimately, representing consumers is a collection of entities that they engage with. In this case, let's say items that they order or click or view has limitations as well because underlying that is an assumption that you can, again, connect these items through some shared metadata. And in my opinion, yes, we are able to achieve quite a bit from ID-based features, for example, and sequence features.
The value that LLMs bring through this profile-based understanding is, again, the semantic nuance of the attributes of the items and the nature of the engagement. So like you mentioned, understanding why a user might be ordering a late night snack or why they might be purchasing a regularly large basket of items could be related to hosting a party or some event of some sort. And that kind of reasoning, again, is not easily achievable through traditional models. They might be learning some of that if there's enough data available from historic behavior. So that's like the second constraint that I think LLMs solve, which is they help us leapfrog that need for those labels. So there's one part that is just a cold start part. There's kind of like an infinite set of intents that you can model, and you definitely don't have the ability to distill that down into any form of taxonomy that you can learn from.
And then the other part is just the fact that this behavior is most likely not already observed in your platform. This might be something that you want your platform to be able to do or provide as a experience to your consumers, but your consumers are not using your platform for that. So that's the feedback loop aspect of it. If you don't have that behavior, you're not going to learn it.
So I think the LLMs really help you break out of that as well.
So it's some kind of using their somewhat internalized knowledge of those LLMs to also be able to generalize, to incorporate different concepts for explainability or as additional means for improving predictions by taking into account this context like, oh, there is a specific holiday, there's probably a party because those models have as foundational background knowledge that Friday evenings people might celebrate more likely than on our usual workday Tuesday evening where they might rather go to do some sports or whatever. And then you have this generalization.
Yeah, quote unquote, as they call it, world knowledge. And I think the trick is figuring out how do you really distill that world knowledge in an effective way into our RecSys and search stack. And this is one way that we've investigated. Can we just let the LLMs do content generation? But we're relying on a very traditional retrieval and ranking architecture after that.
So it's kind of like a natural language distillation of reasoning about the consumer and the intents, but then ultimately representing it in a way that you can use traditional retrieval and ranking stack. So I would say that's one pattern of doing that. And there are, I would say, at least three patterns here. So one is put the LLM in front of your stack, and another pattern is put the LLM as the final stage. That's where I think the concept that you brought up of the fact that LLMs provide a explainability layer as well in many ways.
So for us, that comes from both these memory profiles that we're generating. One major feedback we get from all of our cross-functional partners is now they're able to really think about purchasing patterns, not just in numbers, but from the perspective of the consumer and really bringing a consumer's voice into it. Even though it may be imperfect, it adds the reason it's perceived as that by especially our product and strategy and ops teams is that it's represented in natural language. And a natural extension of that is an area that I am very interested in, I don't think I've seen a very good example of it, honestly, yet, is explaining to consumers why are we showing something to them. Oh yeah, yeah. Big topic. It's good that you bring this up and also leveraging the LLMs for, let's say, embellish around the recommendation or the ranking to explain a bit better why something is displayed to users. And just a former guest, a long time luminary, Joe Konstan said about this, and I quote, sometimes the best way to make things useful is not the recommendation itself, but the information put around it.
And I guess this is where LLMs can definitely help. And honestly, I personally, as a consumer, think that in terms of explainability of recommendations, we have traditionally not been doing our best job. So there's, I guess, still a lot of potential left on the table for this. Yes, indeed. And that's kind of why we also decided to start with the construct of collections.
But again, I wouldn't say this is a perfect user experience either, but let's say on an app like DoorDash, and it's fairly common across e-commerce apps as well, is showing you a collection of items and then putting some verbiage around it and saying like, hey, this is a gluten-free bread aisle for you because we understand that you like gluten-free products. And when you say collection, more concretely, what you refer to as a carousel, for example. A carousel, exactly. A carousel collection, essentially you can think of it as a set of entities, in this case, the items.
It could also be merchants or stores, because it looks like you're improving your home. Maybe you're interested in going to these local furniture shops. Yeah, and doing this can already help the user build trust because they feel understood. At some point, it's also more risky if you don't get it right, possibly, because users can also lose trust because they feel misunderstood.
Yes, yeah, exactly. It's a fine balance, right? You don't want to end up in a territory that loses trust with consumers from, I would say, either by being too creepy, like you don't want it to be too specific, or blatantly wrong. Both extremes are bad.
Again, this is where I'm very genuinely excited about where we stand today in the recommendation space, is the fact that language models give you such flexibility to build these experiences now. A lot of this could have been multiple iterations of product feedback, but now we can actually just prompt LLMs to generate different ideas and test them. And ultimately, with some constraints, I think online testing is the best way to look at what our consumers care about. Previously, what would have taken multiple iterations of product decisions can now be distilled down to just a few aligned objectives and letting builders launch experiments and test. Yep, completely agree there. And maybe something that we will also get to when we talk about shopping agents, because I want to challenge this assumption or this trial quite a bit with reference back to conversational recommender systems. But before we go into that, I want to build the bridge into some topic that I know, since we have a frequent exchange on these topics, excites you very much, which is semantic IDs. But before we dive into semantic IDs, I guess our listeners also deserve kind of a more concrete understanding of what are the actual problems you are responsible with your team to solve. So we have introduced you as the tech lead for search and personalization. And in terms of personalization, well, you can personalize search.
Actually, you should. And when we talk about recommendation, we always, and this, I guess we stressed also in the last episode that I had with Sascha, is we have this locality restrictions that sometimes make the work harder for us, but also sometimes easier. So sometimes we can directly jump into ranking, for example, our stores, as we are naturally constrained to a small in itself set of candidates that we see or that we need to rank and don't need to perform any specific candidate generation. Of course, when it comes to items like products, again, this might look very differently, but when we talk about these different problems like search problems, ranking problems, recommendation problems, could you pick like one very specific problem that could also help us to better understand the work that you are doing and semantic IDs and all the stuff from each of these three spaces that you face, which you think this is important to get right for our consumers so that we can take it as a reference when we talk more about semantic IDs?
I think when we look at where we are building in terms of where and what we're building on, discovery is what we call it, right? So within discovery, there's search, personalization, and now, agentic ordering as well. And these are the common ways that consumers essentially come to our app to ultimately order, but also ideally discover new things. Search, I view as almost as a subset of recommendations nowadays in the sense that search brings in an additional context that general discovery doesn't have, which is the query itself, right? And the query can reveal quite a bit about the intent of a consumer, obviously. And when we talk about what is good in terms of what objectives a search stack should look to solve, I think one of the most important components of that is relevance.
Relevance to the query is supreme for search, right? If you surface irrelevant items for a consumer, it's not the end of the world, especially in an e-commerce app, but it's definitely, again, not building trust for the consumer to come back to that as a discovery surface, right?
So everything needs to be balanced against relevance on search, whereas for general discovery, we can be more creative and we can build in more exploration, obviously.
And through that process, learn a bit more about the consumer beyond the current session and current intent. So that's kind of like how I would differentiate the two more traditional paths that consumers have taken on apps like DoorDash. Then comes this new paradigm of agentic surfaces. So when I say agentic, these are, you know, chatbot style conversational surfaces where the user is trying to achieve a lot of the similar goals as before, which is ultimately, again, ordering items on an app like DoorDash, but it could also be a bit more open-ended in the sense that they don't know yet what store they want to purchase from, for example, or they might be looking for recommendations on the items that may solve the particular problem particular problem that they're looking to address. So it could be like, for example, planning for a party of 10 people, like what should they cook given dietary restrictions of the people, right? And this becomes a very rich surface, which otherwise would have been maybe a multi-click experience, multi-search experience for a consumer can be all condensed down into like, ideally one single ask for the app and like the agent figures out the rest based on what it knows about the consumer, what it knows about the catalog, what it knows about the local fulfillment options and gives a concrete recommendation to the consumer, right? So in all three of these cases, what is quickly becoming apparent for us at least is that one does not replace the other.
I think for a large part of the coming years, these systems will need to coexist, like discovery surfaces have coexisted and developed as parallel surfaces for a while, but like these agentic ordering surfaces or chatbot style surfaces are quickly emerging, but I don't think they will necessarily replace search or discovery, mainly because I think these are still in terms of the user experience requires quite a lot from the consumer, right?
They have to still provide a lot of input to get to where they want to be in terms of like the final cart that they want to order, right? That's one part of it and the other part is novelty and I think it takes time for consumers to adjust to it. So that's one part of it, like the premise is that these three stacks will exist, but even if we say like these evolve to just one pattern of user experience, let's say it's just the chatbot pattern, the underlying infra would still be driven by, I think, a hybrid set of search and recommendation models, right? So what has surfaced as different use cases and also fulfilling different needs of a consumer, because I feel when I search for something, then what I saw in the first place wasn't relevant enough for me to find what I was needing. Or if I come to the platform, you could have already arranged for my grocery shopping basket. I shouldn't put it together myself. And I know that, for example, something that is important to us is to help users build baskets. Whereas we could also say, why do you even need to help me build my basket if you could already come up with a basket yourself?
But then again, we are back to the point, well, we should be very sure about getting it right, because if like half of the items in that basket are wrong, then this also doesn't help you either. So this is like search. So I go to search when I don't find it.
Yeah, exactly. A very concrete example of that is actually just what we call is complementing your cart, right? On DoorDash, we try to surface this on multiple areas of your journey on the app, right? So like the idea is that as you start building your cart and you add items to the cart, we should get even more sure about what are the next set of items that are likely to add, right? And through that, you have some intent prediction of what they might be looking to achieve from this cart. A common example for this is I try to make a salad. So I add some salad, I add mozzarella cheese, and then of course what's missing and what I sometimes forget to order are the tomatoes. And of course, like here we are, there are cherry tomatoes for you to pick Marcel. Yes, exactly. That would be like a perfect situation, right? But again, when we think about it, like traditionally we use like sequence models or like next item prediction, very much limited to the data that we have available on the type of carts that people are shopping for, right?
And then genetic ordering surface helps us break that loop to an extent because we're able to now generate the whole cart given a stated upfront intent by the consumer. But now if you distill that down further, right? Because the premise is that not everyone's going to jump onto chatting with agent to build their cart. You need to now figure out how do you create a learning loop across these stacks so that they don't evolve separately and one gets left behind, right? Because that would be detrimental for the marketplace and also the consumer experience. And I think that's where kind of like the concept and this is not only the case that DoorDash will deliver, I think generally across the industry there's a recognition of the fact that there might be some unification here that we can do. And just going back to your original question of like where I see semantic IDs in this picture, I think this is just one of the substrates for this unified stack, right? Like you need a way to represent the catalog that's learnable by all of these systems at the same time. And if you think of agents, the foundation model involved over there is the language model, right? So ultimately you need something that's an identifier of the catalog that is learnable as a token by these language models and all the research from the Tiger paper, the Plum paper, originally from DeepMind and Google and YouTube really motivated this within the industry. And I think within the e-commerce domain, this is actually even more pertinent, right? Within e-commerce items really carry a lot of meaning. If you look at social media or domains like YouTube, where you're doing video recommendation, yes, the videos themselves have some themes, which you can tie to like some preferences that the user may have, but ultimately there's less friction, right? If you recommend a slightly tangential theme, if you don't get the theme exactly right, people can scroll past that quickly, can ignore it, right? Whereas like in an e-commerce domain, like users expect you to know the inherent meaning of an item. And if you show something that feels completely off, for example, if I showed you in a salad bowl, I showed you like, I don't know, what wouldn't go in a salad bowl? Now I can't think of it. Maybe radishes. Well, actually salad bowls. This is the hardest question of all, like not how to unify certain recommendation, but what does not go into a salad bowl? Nowadays you can put everything in salad.
I will open up the comment section for this for all listeners.
Exactly. I was like, oh yeah, fruits can go in a salad. Oh no, nuts can go in a salad. Now I can't think of anything else. Maybe not a steak.
Yeah. Oh, well. That's rare.
I don't know. Rare, rare.
So not the steak, but the occasion.
It's doable.
Yeah, you can also do steak salad. Okay. So hear us out, listeners, if you can come up with some examples quite quickly. So like, don't ask your favorite chat bot for it. What does not go into a salad? Please post it into the commentary section. I would love to learn now. Okay. Let's say that you're shopping for a party, a birthday party, right? And like, suddenly we're showing you something related to...
He's thinking about sex toys. I know.
And then he was like, I could read your mind because at that very moment it was like, there are also people celebrating possibly that kind of birthday parties.
Exactly. Okay. When we get the idea now, like I'm being really bad at examples here. But yeah, like consumers know when you've gone wrong very well, right? And are, I think, less lenient on a platform like the e-commerce platform.
So non-working examples aside, I guess you already brought up some very important keywords like vast catalog, which is definitely a problem in our domain for the immense amount of products we are facing. You already mentioned like few billion items over and across all merchants.
And that even after the duplication, you can end up with like hundreds of millions. So quite a large corpus that we need to understand, as you said, like in our domain, it's much more important to get the content understanding right. And we have the tools at our hands to do so.
While at the same time, we observe this trend in the industry that is also very nicely reflected by this very first edition of the unified search and recommendation workshop that is co-located with this year's RecSys and where on October 2nd, you will actually also provide a presentation of a paper. And I have it here in front of me. And the name is One Hierarchy, Two Systems, Semantic Product IDs for Discovery, Surface Ranking, and Search Page Query Refirmulation that you wrote together with your colleagues. And here we are, like everybody is talking about semantic IDs, generative retrieval. Of course, transformers are already here quite a long time, but they become kind of their, let's say, model backbone for leveraging these kinds of representations. So Raghav, introduce us. You are the very first one on this show talking about more specifically and practically about semantic IDs. If we take the paper as an example of use cases where you put that to work, when and how did you cross semantic IDs and how did you make them useful for your own work? So I think the concept of semantic IDs is something that I came across through the original papers from Google and YouTube. I initially read it. It actually immediately looked like a research idea that was very applicable to our domain because, again, within e-commerce, we did already have a lot of ways to structure our catalog that can be a bit more learnable. So an example is having structured taxonomies for our item space.
But ultimately, the biggest bottleneck of these approaches were the fact that this was a manual approach, firstly. So we have a taxonomist team that identifies what are categories and how we should split the taxonomy tree and how granular we should get. But there's only so much they can do in terms of granularity because the product use cases are vast. So they have to generalize to the needs of multiple teams. The second is the actual taxonomy itself still doesn't capture the nuances of the specific items. You can call it the coarseness of the taxonomy to an extent, but it's also driven by what you might consider as being arbitrary nodes within a taxonomy.
For example, it may happen that within a taxonomy, we've classified as terminal nodes all the variants of sourdough bread, but there's an equivalent taxonomy for white bread as well, that is not sourdough, but trying to capture everything that's not sourdough in one larger node. That's not ideal because, again, this goes back to the question of how do you translate that into models that are going to retrieve and rank very specific items for the intents of the consumer.
So this was already a motivating factor to have something like semantic ID available to us as a primitive for our models. Because for those who are not aware of the nature of semantic IDs and how they're built, a high-level review is that semantic IDs use embeddings. Those embeddings can be generated from the content of the items themselves. So these could be multimodal semantic image, video-based embeddings that capture the meaning of the entity itself, or they can also be learned embeddings from, for example, collaborative filtering models, right?
Two-tar based embeddings, for example. And the idea is that a combination of these or one of these input embeddings would be used to then do a recursive clustering such that you're able to generate codebooks at each stage. And this means that, for example, today at DoorDash we use what is referred to as residual k-means. We cluster embeddings recursively, as I mentioned, using k-means as a clustering algorithm. And at each stage, we're essentially generating n clusters, taking the residual of each cluster for the next stage and then clustering that residual over time, and then ultimately doing so so that the represented sequence of IDs try to form natural node of similar items based on the content or the collaborative signals that are learned from the original embedding. But a natural kind of unique property that emerges from this is the fact that these IDs are hierarchical in nature. So they essentially form a taxonomy, a learned taxonomy. Now, based on our investigations and research in this area, these taxonomies may not be the best to explain our catalog to a human, right? A human may not naturally cluster these items or structure this taxonomy in the way that semantic IDs do, but they do cluster on meaningful attributes of items that would be non-obvious to, again, a human building of a very structured taxonomy. Not obvious, but understandable, because I assume that you evaluated this by looking at the clustering, which then might seem odd at first, because you say this would not be like your or a different person's way of clustering them.
But if you think about it, you seem to understand why this cluster has been chosen or that it actually makes sense or that it clusters to some certain aspect and that the clustering itself is somewhat like humanly interpretable or understandable. Correct. Yeah. So the way we evaluate these semantic IDs is both through quantitative measures and also qualitative measures. And the qualitative measures are largely looking at things, common terms and if-if style encodings of the attributes of the items themselves and looking at common words that these items share at a given layer. So like, you know, a taxonomist may, for example, in the clothing line, decide that, you know, jeans is to be split based on the type of material, then by what kind of design it is. So, you know, whether it's torn jeans or bleached jeans, etc., and then maybe by gender. Semantic IDs would kind of like cluster maybe by combination of different attributes, right? So it could be like torns, men jeans or bleached mom jeans, right? And this is like very fascinating to us because some of these clusters don't even exist in our taxonomy today, right? This goes back to the coarseness problem. So we're not able to get down to that fine level detail of our items from purely our taxonomy. And then secondly, it's learning a lot of this without the metadata available in our catalog, right? So what we're using today at DoorDash is our largely content-based embeddings. So we passed in all the rich metadata that we have, but there, it's not structured. And quite often the metadata is incomplete, but it's the embedding model, which we're using a large language model for this as well, to generate these embeddings, inherently capture some world knowledge about these items, and then use that in the clustering process itself. Okay. So when you say like metadata, you mostly focus on text data at the moment?
Yes, because images are fairly noisy, especially on a platform like ours, different merchants have different ways of photographing items. So image data we found to be fairly noisy for this use, but we use all the text data that we have available today. But most of the time this boils down to the item name and to some extent the item description. Okay. In the paper, you mentioned, like as you said, the item name, brand, size, and then you create this kind of textual profile from this. Correct. Yes. So from there, we're finding that semantic IDs are able to cluster these in a pretty unique way to our taxonomy. So the paper actually looks at two use cases over here. And our goal was to really identify some easy use cases that can prove that semantic IDs as a primitive can work for our goal of having some unified understanding of our catalog across search and personalization. So the simplest use case we could think of to test this on personalization was to test this as a feature in our ranking models. So our ranking model today is a deep learning-based multitask, multi-label model. So we're trying to optimize on all the engagement objectives on our platform. So that's click-through rate, add to cart rate, conversion rates. And we have pretty rich representations of items and consumers through existing IDs, taxonomy IDs, item IDs, and then also a lot of dense features, aggregate features on taxonomies, on items, on stores.
And these could be like how frequently have consumers purchased an item from a given store, as granular as that, to how often have consumers purchased that category of items. And our taxonomy, for example, today is divided across four levels. So the fourth level being the most granular level, and we have features that span all of those levels. So it's a fairly robust item ranker.
And our goal was to see, okay, it was not only to see if we can add semantic IDs as features, and if that improves the model, but actually can we simplify this model and remove features from our model and still achieve improvement in online metrics? And that's exactly what we saw.
So we use semantic IDs. In this case, we actually largely focused on, can we prove the value of semantic IDs purely from dense features? So can we aggregate engagement across semantic ID layers instead of taxonomy layers? And do they provide better input features than the taxonomy or item based features? And we saw exactly that. Just to get this right, how this works when you say dense features based on these layers. So from the papers, chapter three, I can see that you chose three stages. So every item's semantic ID is represented by a tuple of three integers to say so, which then stand for the corresponding clusters on the corresponding level.
And when you say you learn dense features, for example, the past four, eight, 12-week purchase count, then you learn this in a cross feature between the user and the corresponding semantic ID level. For example, let's say it would be fruit, berries, and then you have strawberries. So that you basically say the user ordered three times these various strawberries in the past. So that means on the fruit category, you have at least a count of three, at least, because they might have purchased other different sorts of fruit. Then in the berries category, you also have a three at least. And in the strawberries, you definitely have three.
Is this how I could think about this as a concrete feature, for example?
Correct. That's right. And so just to clarify a couple of things. So the semantic IDs that we generated, that's described in the paper, yes, they're three-level semantic IDs. But the features that we ended up using, even in these aggregate dense features, were based on semantic IDs, like the n-grams of semantic IDs. So our semantic ID codebook size is like 512 cubed today. So each level has 512 possible IDs. And in order to scale that effectively, and also build dense features that are meaningful, that are not too sparse, effectively, we went down the route of actually building n-gram tokens, and then using those n-gram tokens to aggregate our engagement features at the item and consumer level. So yeah, exactly like you said, it could be something like an n-gram that represents all the way down to strawberries, like a strawberry cluster of items based on the semantic IDs. And then based on the engagement, at the item level, we would look at the engagement across the surfaces, and also at the consumer item level engagement. So just those features by themselves were impactful for our model, actually added much more values such that we could even, I think we ended up removing around 17 to 20 features from our original model that were taxonomy based. What does it mean, comparatively?
Comparatively, so today we have around 250 features in the ranker. This is a combination of dense features and sparse features. Dense features are, I would say, in the range of 180 or so features. So we were able to effectively remove 20 of our top.
So we basically looked at the feature importance of our top dense features, and then how many of those could we replace from the top 50. So we were able to replace 20 of the top 50 dense features with these semantic ID based features in our ranker, and still achieve a pretty sizable lift in our online metrics with this model. So that was one indication of semantic IDs, even though they're content based today for our construct, they work well for personalization tasks like recommendations. The other use case that we wanted to test, which was to see how well they work for our search, was using semantic IDs for query reformulation. So historically, on our app, this goes back to the idea of consumers building a cart, and most of the time, there's like a complementary nature to the items that they build or are present in a cart.
That same idea can be translated to search, but it's a little broader. So users could be searching for berries, and then one path they end up taking could be that they are searching for berries and then refining their queries over multiple terms to get to the exact type of berry that they want, organic strawberries, for example, or organic Drisorolos strawberries, like a specific brand.
The other path they could take is they search for berries, they added something, and then now they want a complementary item for the fruit salad that they're building, going back to salads.
But the idea is that query reformulation becomes a natural friction for consumers. You want to provide recommendations of the next query that they could be searching for. Historically, this is kind of like a graph traversal problem, looking at historical queries and learning from that. But it's pretty hard to learn from pure string queries, generally. Our goal was to see, can we represent queries as semantic IDs and use that to power this query reformulation layer.
The underlying that was the idea that semantic IDs will represent some granularity of the item attributes that would be equivalent to the first path, like getting to more refined search terms to get to a specific item. So that's kind of like traversing the hierarchy and going to a deeper node within a specific semantic ID path. And the other idea was that they might want to look for substitutes or complementary items as the next path. And that's kind of, again, a tree traversal problem in many ways, trying to find other nodes that you can traverse instead of going deeper and doing a breadth-first search for the queries that you can recommend to the consumer for that intent. So that was also actually a pretty good evaluation in our case of how well we are able to use semantic IDs for search. We saw that engagement with query reformulation with semantic IDs as a representation there did increase the click-through rate cross-search. I would say this is probably the first small scale test for us to really prove its value on search, but this gives a good enough confidence to us that semantic IDs can be that meaningful primitive that we use to start building this stack across these three surfaces, search, discovery, and the agentic ordering surface with a shared kind of content understanding and catalog understanding.
Stay Forever So this means that from now onwards, we can definitely expect you to double down on the usage of semantic IDs given this first initial proof of business impact by using semantic IDs for two different use cases. But in the end, use cases that not yet share a complete same backbone, but share some same representation for items here, which is mainly the semantic IDs that are used for both use cases here.
Vint Cerf Correct. Yeah. And I think that brings up an interesting point, right? Like, and this is something that we've been thinking of as well as to what does it mean to have a unified stack truly, right? I would say today, like what we're trading towards is the fact that we have a common substrate. So semantic IDs are one example. Memory or consumer profiles that are understood by language models as inputs is another, right?
There are a few others that we can talk about, but these are like the two common blocks, let's say. On top of that, one concept that is becoming evident again from industry is that there is a concept of like a foundation model, right? Like there's some reusable backbone potentially, or even a reusable model that learns across surfaces and provides useful representations that may not necessarily be the models that power a given surface, but provide inputs to that surface, right? And then you have like task specific models or surface specific models that use these two sources, like these primitives and the reusable backbones to bootstrap themselves for the objectives of the business and the surface. So that's kind of like the architecture, I think we're heading towards. There is also obviously like a nature of this problem where we say, you know, we just have a unified everything. So like a single massive, maybe one single model that does retrieval ranking all together across search, personalization and the agent.
We wanted to talk about this move from generative recommendations to shopping agents, which for me it doesn't necessarily seem needs to be a move, but maybe it could be two things alongside each other. So we are advancing the field of recommendations and search, especially with unified architectures, with semantic idea representations, with foundation models, et cetera, that we use for these use cases, but just make them, let's say more comprehensible and using better representations that are then joined with user behavior to provide better or more relevant, but also more engaging, more discoverable output for users. While at the same time, we have this move into agentic and you also brought it up, we have this product, ask door dash. I'm slightly and wearing my show host hat here, slightly skeptic that this is really flying. I mean, nowadays you go to almost every online shop, e-commerce website. It's not just only constrained to e-commerce, but you may find it as well in streaming media. And this whole topic feels a bit like resurfacing conversational recommender systems. I've never been a big fan of them, but nevertheless, worth discussing and challenging because I feel like recommendation should leverage all the information the system, the platform has about me, about my behavior that are reflective of my preferences, leveraging the content and its representation and so on and so forth. And when I come there, they should use this data associated with predicting my next best action, putting things into a proper ranking, learning from others and from my profile and behavior to already show me what's relevant for me in an ideal or perfect world. I don't want to go there and having to go through the burden of explicitly writing down each and everything. In a perfect case, the system is anticipating what I want and I don't need to articulate what I want. So is the existence and their popularity of agents more a reflection that our recommendation systems and search systems are still not standing up to this expectation properly? Let's put it differently. If recommendation and search would work properly and let's put recommendation systems primarily there, would agentic systems be necessary? I think it's an interesting question, but the way I view it is that it's not that agentic systems themselves are a new system, but rather a user experience question.
And this is where I think the stack itself, technically what powers this, whatever it may be, whatever nature it may be. So when we say search, it's usually like some search results page with some search bar on top, some items at the bottom. It's a user experience, right? It's like an expectation now of consumers to see a similar representation of the results across all apps, wherever they're searched. Same thing with recommendation and discovery. There's an expectation that there's some landing page, like a homepage, for example. And then there's just a bunch of stuff thrown at users and they just, this is like a natural expectation. You would never feel odd that an app is doing that. And I think agents and these conversational agents in particular are in many ways a similar new experience that I would say to an extent, just through the proliferation of chat GPT and Claude and all of these consumer-facing chatbots, is also becoming a norm, just like a well understood concept that people expect when going to apps. And now the question becomes, from a user experience point of view, what is working and what is not working? And where are the advantages and whether even this form makes sense?
That's one part, right? And then the other part is like, are our ML approaches, are our technical approaches meeting the requirements also of this evolving user experience? So I think from the user experience standpoint, I agree with you. I don't think most people want to type out what they expect, what they want to order, for example, from an e-commerce platform. I don't think people generally want to type. Yes, this was my point. People are lazy. People are lazy, right? And they don't want to be descriptive. Oftentimes it's also too expensive on the mind to think through all of the ways you want to describe what you want to do. Now that might be a question of whether the medium is wrong. Maybe typing is wrong. Maybe voice narration is better. Maybe people are better just voicing out their needs. Well, we directly jump ahead to their human, to the brain computer interface. Exactly, right? I'm just sitting here and then there comes my doordash order because I was like, oh, I forgot to order my lunch. We will have to see how that evolves. But one thing is clear to me, to your question, I don't think our current evolution of user experiences of how we show search results and how we show recommendations has evolved recently. All apps look the same.
As you say this, and for me this leads to the point, it's not an either or, but it's two different ways of interacting with an offering, two different ways of navigating and getting a problem solved in the end for the user. I would say they're two different ways, but neither are optimal. There might be a situation where they collapse into one. So might this be a bit like the iPhone moment for personalized experiences? Because nobody before the iPhone existed needed a smartphone. Nowadays, everybody has a smartphone. So is this similar? So just the fact that we are well aware of how search and recommendation work and what we are going to expect, this is something that we think, oh, this is sufficient. And now comes somebody with a totally new product and that creates its own demand because it shows what is possible and people then seeing what is possible start to have that need for it. But I think the intermediate or the first step over here would have to be, hey, we can't have these three different ways that consumers are expected to engage with the app, to get to what they want. It should just feel on an experience level as well, a unified approach. So we can unify the stack underneath it as much as possible. And that is largely driven by the fact that it helps iterations, it helps a velocity of launches and testing new features. And also from an ML perspective, it goes all the way back to these multitask models, mature expert models, foundation models. The more data from more surfaces that you can train the model with, that's the benefit that you get technically. But that I think is a very neat and nicely, I think a problem space that is very nicely tied together. I think the tricky part is what does that mean from a unified user experience? One way I think about it is what if an app was like a living entity? And it didn't feel like a multi-turn agent, but it just had a more fluidity in what it's presenting to you. So instead of you going to a search page to search, to go to a homepage to discover other local stores, it's just that it just understands based on some sort of input, it could be voice input, conversational in some nature, or through the way you're scrolling, it evolves the user experience into a search flow or into chatbot flow or a discovery flow. Our surfaces, our user experiences are not malleable enough right now to truly be unified. So I think that is an ideal state, in my opinion. I just have to open an app and it just evolves the user experience without me needing to tap three different icons to get to where I want to be. Yeah, this is, I guess, also a matter as it is described quite a lot among practitioners and researchers that we sometimes too narrowly think about recommendation and search and personalization more broadly as an solely algorithmic problem, but it's more like a system problem. And a system also involves like UI, UX, like how do users navigate content? How do they get through it and go to the point and not just, oh, I tweak my algorithm by an additional 5% uplift in MIR or NDCG or whatever. And now here we are, but it doesn't translate into a real impact because the flow is still the same and presenting the bottleneck or the bottleneck, like the user flow and not like the excellence of the actual offline results of an algorithm. But I guess all of that might be also the topic of this very workshop where you, my dear listeners, have the chance to meet Raghav in person and discuss even more about this and other related topics with him. But this won't be the only occurrence that people have the chance meeting not just you, but also me and Pawel as the Doordash representatives at this year's RecSys because on the very first day of this year's Recommender Systems Conference, we will also be giving a tutorial, as I announced at the very beginning of this episode.
So a tutorial where we actually want to share all the solutions, challenges, and learnings from doing personalization in our domain and more concretely recommender systems. And part of that will also be Raghav sharing about semantic IDs, about specific challenges in this domain.
But Raghav, share a bit more on what people can expect from our tutorial at RecSys on September 28th at 8am, 9am in the very morning. Yes, the very first tutorial to kick RecSys off.
That would be pretty fun. Yeah, I'm excited to meet a bunch of folks at that session and beyond.
I think the core premise of the reason maybe, Marcel, that we decided to give this tutorial is just the fact that, as I alluded to before, multi-sided marketplaces are becoming very prevalent. And it's kind of a recent phenomenon in many ways. E-commerce has existed for a while, but multi-sided marketplaces provide a different set of problems and challenges. And when Marcel and I started talking about potentially doing this tutorial, that was exactly what we identified as being a gap that has not been previously covered as far as what we saw at RecSys and other recommendation conferences. And that's essentially the structure of the tutorial that we are covering. We are going to be talking about the product and business aspects of running a multi-sided marketplace and then tying that back to how you would structure your recommendation and personalization models to really benefit such a marketplace. So, yes, we'll cover a lot of learnings that we've had building retrieval and ranking models, doing multi-objective optimization, doing exploration, and through that, hopefully provide some insights to other practitioners who may be operating in the same space or in adjacent spaces. So that's definitely one goal.
The other goal is to, again, as you mentioned, really talk about not only semantic IDs, but the generative Rexx space and what does it mean to build in this new kind of emerging area. We've had a few examples that we've shared through papers so far and blog posts, but there are other examples that we've not yet published in this space. So we will cover some of our learnings over there as well. And ultimately, the goal of a tutorial is obviously to engage with you all who hopefully will attend the tutorial as well. And through that engagement, through the questions that you might have, we also hope to learn and essentially use that as an opportunity for us to brainstorm future ideas. So very excited. It'll be a pretty fun RecSys, I think. So excited to see everyone over there. Yeah, it's going to be cool. And it's going to be the homecoming for the RecSys, which started out in 2007 in Minneapolis and is now coming back to the very same city.
Even though I have to admit, I'm slightly even more excited about the venue of next year's RecSys, which will actually be Hawaii. So this is definitely going to be special and exciting as well. But for the moment, we are looking forward to meeting the community towards the end of this month in Minneapolis, in Minnesota, in the US this way. Yeah, getting to the end of this episode, maybe there's one question left that I would be curious about. You're now working in this field for more than five years at Doordish. It's a fast moving company, very exciting work, can be also stressful at times. What would be your advice for current practitioners who are, for example, entering the field, navigating the space, also given all the tools that we have nowadays from Clodcode to Codex to using Gemini assistance and so on. So I feel sometimes like it's very hard also for people who have collected some experience already to navigate this quickly moving space. What would be your advice for people to be successful or to find their way in this even faster moving world that we are operating in nowadays? I don't think I have a proven answer or a working answer to this yet. What's really fascinating is actually just going back to my move from chemical engineering through the world of energy policy and then ultimately ending up in tech and working on search and recommendations. What has always driven me has been probably a natural curiosity and then the fact that all of my previous experiences did not provide the pace at which my curiosity moves. And now looking back at it, it's proven to be a double-edged sword because the pace has gone out of hand in many ways.
So the tech industry has always moved really fast. Especially within the quote-unquote ML space and now the AI space, techniques are rapidly evolving. Previously, it might have been evolving, let's say, on a six-month to a yearly basis. Now it's like at least every quarter, not every month, there are significant new pieces of research or model releases.
And I'm not only talking about language models, but even within the information retrieval literature, REX's literature, there's so much change that's happening. So I don't think it's that new in many ways. I think the rapidness of change is definitely increasing, but the fact that change is happening is not new. And I think that is probably the way I look at it is that does attract a certain personality for sure. So you do have to be truthful to yourself whether this is really your true calling. And the other part of it is that even if you do have that natural curiosity and the interest to quickly learn new things, you have to balance that with the fact that we are all human. So being burnt out is fairly easy in such a space.
But I think to an extent, we can leverage a lot of the recent advancements in language models as these knowledge building tools to our advantage as well. And that's what I've been trying to do as much as possible. Whilst I don't have multiple agents running the background for me, like other people claim, I do try to use it to rapidly learn about an area. Also, what's very important is to understand what you should not waste your time on. Because as you have rapid evolution, there are going to be a lot of dead ends. That's, I think, equally important for any practitioner to develop a skill on. And I wouldn't say I'm perfect at it. But what I definitely try to do is I leverage these Clod or ChatGPT to quickly filter out content that I don't need to focus on. I think that is definitely a skill to have and learn. But I actually associate that in many ways to the same thing that we did when Google search was becoming big. And I might be dating myself a bit, but I remember my parents telling me, like, hey, you're relying too much, or my teachers telling me you're relying too much on Wikipedia and Google. Lo and behold, that is now very much the norm and probably outdated in many ways. So I think we all have to adapt. And I think in many ways, the tech industry provides that opportunity, unlike other spaces that I've been in. But yeah, the ability to do that in a healthy way is still something that I am also learning, to be honest. So probably no answer. No, I guess this is far from being no answer. This is definitely great advice. So spot on there. There were lots of things that also resonated with me, which I think is definitely useful. And also, as you said, not knowing what to pay too much attention to. No, I think that people can definitely leverage this. I hope so.
Cool. Maybe last question for this interview, and this might not come as a big surprise, or what would be the person that you would like to know more about in terms of their RECSPERTSise, RecSys search personalization work on this podcast?
Oh, that is an interesting one. So I've been seeing a lot of pretty amazing work coming out of Spotify recently as well. And Netflix also has been a very predominant contributor in this space.
So I would think someone who has worked across these two companies and seeing the evolution of recommendations, because I think Spotify is doing pretty interesting work in the generative recommendation space. And Netflix basically set up the foundation of how we understand recommendations in a large way. I don't know if I can think of a name. Maybe I might be butchering the name, honestly. It's E.S. Reimond. Ah, Yves Reimond.
Yves Reimond, yes. That would be an incredible guest to have. I feel like he basically has done that. That sounds great. Okay. This is basically a person who has seen both companies from the inside. So E. Reimond, we need to talk. But I'm also looking forward to meeting you in the U.S.
So hear us out and let's talk to continue the discussion about the topics that we also touched in today's episode, which for the moment we will conclude. I'm glad that you were able to share so many details about the work and also the results, the practical results of this work that you were able to harvest and to see and also like as a good proof to carry on the work that you are doing on generative recommendations. So Raghav, I'm not just happy to have you as my colleague and as a person I'm having the opportunity collaborating with, but first and foremost for today to have you on this podcast, honor experts and sharing your expertise. Thank you very much for that. Thank you Marcel for having me here as well. It was a fun conversation. Appreciate it. Cool. Then I guess there's not so many weeks left, about four of them until we also meet again in person.
Again in Minneapolis for RecSys and we kick off the very RecSys with our tutorial together with Pawel.
And for the moment, thank you again and see you at RecSys. See you all at RecSys. Bye.
Thank you so much for listening to this episode of RECSPERTSs, recommender systems experts, the podcast that brings you the experts in recommender systems. If you enjoy this podcast, please subscribe to it on your favorite podcast player and please share it with anybody you think might benefit from it. If you have questions or recommendation for an interesting expert you want to have in my show or any other suggestions, drop me a message on twitter or send me an email.
Thank you again for listening and sharing and make sure not to miss the next episode because people who listen to this also listen to the next episode. Goodbye.
So
