<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Redis on 0AndWild_log</title><link>https://0andwild.com/en/tags/redis/</link><description>Recent content in Redis on 0AndWild_log</description><generator>Hugo -- gohugo.io</generator><language>en-US</language><lastBuildDate>Fri, 17 Jul 2026 16:20:00 +0900</lastBuildDate><atom:link href="https://0andwild.com/en/tags/redis/index.xml" rel="self" type="application/rss+xml"/><item><title>Designing a Product Ranking System</title><link>https://0andwild.com/en/posts/260717_ranking_system_design/</link><pubDate>Fri, 17 Jul 2026 16:20:00 +0900</pubDate><guid>https://0andwild.com/en/posts/260717_ranking_system_design/</guid><description>&lt;img src="https://0andwild.com/" alt="Featured image of post Designing a Product Ranking System" /&gt;&lt;h2 id="tldr"&gt;&lt;a href="#tldr" class="header-anchor"&gt;&lt;/a&gt;TL;DR&#10;&lt;/h2&gt;&lt;p&gt;I designed a daily product ranking system based on user behavior events.&lt;/p&gt;&#10;&lt;p&gt;My first approach published product views, likes, and successful payments to Kafka. &lt;code&gt;commerce-streamer&lt;/code&gt; consumed the events and updated ranking scores in a Redis Sorted Set in real time. Redis was a good fit as a serving store because it supports fast Top N queries and rank lookups for individual products.&lt;/p&gt;&#10;&lt;p&gt;As I worked through the design, though, I began to question whether Redis should be the ranking system&amp;rsquo;s only Source of Truth. Expired or lost data is difficult to recover, and changing weights requires historical metrics if we want to recalculate past scores. Supporting hourly, weekly, and monthly rankings makes retaining those source metrics even more important.&lt;/p&gt;&#10;&lt;p&gt;This post starts with the Redis design, then explores an alternative that stores source metrics in an RDB and uses Redis to serve only the Top N results needed for queries.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="my-first-reading-of-the-requirements"&gt;&lt;a href="#my-first-reading-of-the-requirements" class="header-anchor"&gt;&lt;/a&gt;My first reading of the requirements&#10;&lt;/h2&gt;&lt;p&gt;Initially, ranking sounded like a simple matter of calculating a score for each product and sorting the results.&lt;/p&gt;&#10;&lt;p&gt;On closer inspection, the more important question was how to interpret different user actions.&lt;/p&gt;&#10;&lt;p&gt;A product detail view signals interest, but less strongly than a purchase. A like is stronger than a view, but doesn&amp;rsquo;t necessarily lead to revenue. A successful payment is the strongest signal, yet using the sales amount directly could let expensive products dominate the ranking.&lt;/p&gt;&#10;&lt;p&gt;I used the following signals:&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Event&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Meaning&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Ranking effect&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Product detail view&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Mild interest&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Increase view score&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Like&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Explicit interest&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Increase like score&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Unlike&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Withdrawn interest&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Decrease like score&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Successful payment&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Purchase conversion&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Increase sales score&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;Views, likes, and sales use different units, so I kept their metrics separate and applied weights when calculating the final score.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;score = carry&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; + viewCount * viewWeight&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; + likeCount * likeWeight&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; + ln(1 + salesAmount) * salesWeight&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The sales contribution is based on &lt;code&gt;price * quantity&lt;/code&gt;, with a logarithm applied. Using the raw amount could give expensive or already popular products an excessive, persistent advantage once they reach the top.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="first-design-real-time-redis-rankings"&gt;&lt;a href="#first-design-real-time-redis-rankings" class="header-anchor"&gt;&lt;/a&gt;First design: real-time Redis rankings&#10;&lt;/h2&gt;&lt;p&gt;My first design used Redis as the real-time ranking store.&lt;/p&gt;&#10;&lt;figure class="mx-auto"&gt;&lt;img src="https://0andwild.com/posts/260717_ranking_system_design/current-ranking-architecture.png"&#10;&#9;&#9;&#9;alt="Initial design for daily product rankings using Redis" width="1100"&gt;&#10;&lt;/figure&gt;&#10;&#10;&lt;p&gt;The flow is roughly:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;A user views a product, likes it, or pays for an order.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;commerce-api&lt;/code&gt; records a domain event.&lt;/li&gt;&#10;&lt;li&gt;The event is stored in a transaction outbox.&lt;/li&gt;&#10;&lt;li&gt;An outbox relay publishes it to a Kafka topic.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;commerce-streamer&lt;/code&gt; consumes the event.&lt;/li&gt;&#10;&lt;li&gt;It updates the metrics and final score in the daily Redis ranking keys according to the event type.&lt;/li&gt;&#10;&lt;li&gt;The ranking API reads ranks and scores from Redis, adds product information from MySQL, and returns the response.&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;The outbox addresses the gap between the API transaction and Kafka publication. If an order or like change commits but its event fails to publish, ranking data can drift from the actual state. Recording the outbox row in the same transaction as the domain change lets a relay publish it separately.&lt;/p&gt;&#10;&lt;h2 id="why-use-date-specific-redis-keys"&gt;&lt;a href="#why-use-date-specific-redis-keys" class="header-anchor"&gt;&lt;/a&gt;Why use date-specific Redis keys?&#10;&lt;/h2&gt;&lt;p&gt;These are daily rankings, so the keys include a date.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:metric:view:{yyyyMMdd}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:metric:like:{yyyyMMdd}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:metric:sales:{yyyyMMdd}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:metric:raw-sales-amount:{yyyyMMdd}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:metric:carry:{yyyyMMdd}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:all:{yyyyMMdd}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:processed:{yyyyMMdd}&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;I separated the metric keys because weights can change.&lt;/p&gt;&#10;&lt;p&gt;If we store only the final score, changing &lt;code&gt;viewWeight&lt;/code&gt; or &lt;code&gt;likeWeight&lt;/code&gt; leaves us without enough information to reinterpret the old value. Separate view, like, and sales metrics let us recalculate &lt;code&gt;ranking:all:{date}&lt;/code&gt; using the current weights.&lt;/p&gt;&#10;&lt;p&gt;&lt;code&gt;ranking:processed:{date}&lt;/code&gt; tracks duplicate events. Kafka consumers may process messages at least once, which means the same event can arrive again. If an &lt;code&gt;eventId&lt;/code&gt; has already been processed, it must not increase the score a second time.&lt;/p&gt;&#10;&lt;h2 id="what-worked-well-about-redis"&gt;&lt;a href="#what-worked-well-about-redis" class="header-anchor"&gt;&lt;/a&gt;What worked well about Redis&#10;&lt;/h2&gt;&lt;p&gt;A Redis Sorted Set felt like a natural match for ranking: it stores members with scores and supports fast retrieval in score order.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ZINCRBY ranking:metric:view:20260717 1 productId&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ZREVRANGE ranking:all:20260717 0 19 WITHSCORES&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ZREVRANK ranking:all:20260717 productId&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Redis commands cover both Top N queries and rank lookups for a specific product.&lt;/p&gt;&#10;&lt;p&gt;The API doesn&amp;rsquo;t need to run expensive aggregations. It reads ranks and scores from Redis, then retrieves display information such as product and brand names from MySQL.&lt;/p&gt;&#10;&lt;p&gt;Another benefit is freshness. Scores change as soon as events are consumed, so user actions are reflected quickly. A batch-only design introduces at least the scheduling delay; here, Kafka consumer lag accounts for most of the delay.&lt;/p&gt;&#10;&lt;h2 id="end-of-day-carry-over"&gt;&lt;a href="#end-of-day-carry-over" class="header-anchor"&gt;&lt;/a&gt;End-of-day carry-over&#10;&lt;/h2&gt;&lt;p&gt;Starting every product at zero each day would feel abrupt. Products that were popular yesterday shouldn&amp;rsquo;t necessarily disappear at midnight.&lt;/p&gt;&#10;&lt;p&gt;I planned a carry-over at 23:50 each day, using that day&amp;rsquo;s Top 100 products to seed the next day&amp;rsquo;s scores.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Read the Top 100 from ranking:all:{D}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Multiply each final score by 0.1&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Apply to ranking:metric:carry:{D+1}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Also apply to ranking:all:{D+1}&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The Top 100 limit controls Redis memory usage. Carrying every product forward would become expensive as product counts and daily keys accumulate.&lt;/p&gt;&#10;&lt;p&gt;Keys also have a TTL. Together, expiry and the carry-over limit keep Redis from becoming an indefinite historical store.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="limitations-of-the-first-design"&gt;&lt;a href="#limitations-of-the-first-design" class="header-anchor"&gt;&lt;/a&gt;Limitations of the first design&#10;&lt;/h2&gt;&lt;p&gt;The Redis design is simple and fast for daily real-time rankings. Looking at it from an operational perspective reveals several limitations.&lt;/p&gt;&#10;&lt;h2 id="1-can-redis-be-the-source-of-truth"&gt;&lt;a href="#1-can-redis-be-the-source-of-truth" class="header-anchor"&gt;&lt;/a&gt;1. Can Redis be the Source of Truth?&#10;&lt;/h2&gt;&lt;p&gt;This was my biggest concern.&lt;/p&gt;&#10;&lt;p&gt;Redis works well for serving ranking data, but this design needs more support for preserving the source information. TTL expiry removes data. After an outage or operational mistake, we need enough evidence to reconstruct rankings, and a Redis-centered design alone doesn&amp;rsquo;t provide that.&lt;/p&gt;&#10;&lt;p&gt;Retaining Kafka topics and replaying events is an option. But replaying all events just to recalculate one historical period can be costly. If the event schema or consumer logic has changed, reproducing the original result may also be difficult.&lt;/p&gt;&#10;&lt;h2 id="2-changing-weights-and-recalculating-scores"&gt;&lt;a href="#2-changing-weights-and-recalculating-scores" class="header-anchor"&gt;&lt;/a&gt;2. Changing weights and recalculating scores&#10;&lt;/h2&gt;&lt;p&gt;The first design&amp;rsquo;s separate metric ZSETs allow recalculation while the daily data remains in Redis.&lt;/p&gt;&#10;&lt;p&gt;After the TTL expires, that option disappears. A request a month later to recalculate last week&amp;rsquo;s rankings with today&amp;rsquo;s weights cannot rely on the remaining Redis data alone.&lt;/p&gt;&#10;&lt;p&gt;Runtime weight changes make the distinction between metrics and scores important. A score is the result of a policy; the metrics describe what happened.&lt;/p&gt;&#10;&lt;h2 id="3-rankings-across-different-periods"&gt;&lt;a href="#3-rankings-across-different-periods" class="header-anchor"&gt;&lt;/a&gt;3. Rankings across different periods&#10;&lt;/h2&gt;&lt;p&gt;Date-based keys look sufficient for daily rankings. Hourly, weekly, monthly, and yearly rankings make the key strategy more complicated.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:all:daily:20260717&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:all:hourly:2026071713&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:all:weekly:2026W29&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:all:monthly:202607&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;We could update every period&amp;rsquo;s keys when consuming each event. But then a single event modifies multiple keys, and weight changes require recalculating more sets of scores.&lt;/p&gt;&#10;&lt;p&gt;At that point, retaining source metrics elsewhere and publishing each period&amp;rsquo;s Top N to Redis becomes a more natural design.&lt;/p&gt;&#10;&lt;h2 id="4-should-every-product-live-in-redis"&gt;&lt;a href="#4-should-every-product-live-in-redis" class="header-anchor"&gt;&lt;/a&gt;4. Should every product live in Redis?&#10;&lt;/h2&gt;&lt;p&gt;The memory question also becomes more important as the catalog grows.&lt;/p&gt;&#10;&lt;p&gt;Storing 100,000 products is something we can test readily. At one million or ten million products, multiplied across daily keys, Redis memory cost becomes harder to ignore.&lt;/p&gt;&#10;&lt;p&gt;Most ranking API requests need only the Top N. Keeping those results in Redis and the complete metrics elsewhere gives each store a clearer role.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="revised-design-rdb-metric-sot--redis-top-n"&gt;&lt;a href="#revised-design-rdb-metric-sot--redis-top-n" class="header-anchor"&gt;&lt;/a&gt;Revised design: RDB Metric SOT + Redis Top N&#10;&lt;/h2&gt;&lt;p&gt;To address those limits, I explored storing source metrics in an RDB and serving only query-ready Top N results from Redis.&lt;/p&gt;&#10;&lt;figure class="mx-auto"&gt;&lt;img src="https://0andwild.com/posts/260717_ranking_system_design/rdb-metric-sot-architecture.png"&#10;&#9;&#9;&#9;alt="RDB Metric SOT and Redis Top N ranking architecture" width="1100"&gt;&#10;&lt;/figure&gt;&#10;&#10;&lt;p&gt;Redis changes roles in this design. In the first version, it was effectively the Source of Truth for real-time rankings. Here, it is a read model for fast queries.&lt;/p&gt;&#10;&lt;p&gt;Source metrics live in an RDB such as MySQL.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking_daily_metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- metric_date&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- product_id&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- view_count&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- like_count&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- sales_amount&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Weights are deliberately left out of those records. Views, likes, and sales amounts describe recorded activity; weights are policy, and policy can change. Storing metrics before weights are applied makes recalculation easier.&lt;/p&gt;&#10;&lt;h2 id="event-consumption"&gt;&lt;a href="#event-consumption" class="header-anchor"&gt;&lt;/a&gt;Event consumption&#10;&lt;/h2&gt;&lt;p&gt;The event pipeline stays largely the same. User actions reach Kafka, and &lt;code&gt;RankingMetricConsumer&lt;/code&gt; consumes them. Instead of incrementing Redis scores directly, it upserts metric rows in the RDB.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;PRODUCT_VIEWED -&amp;gt; view_count + 1&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;PRODUCT_LIKED -&amp;gt; like_count + 1&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;PRODUCT_UNLIKED -&amp;gt; like_count - 1&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;PAYMENT_SUCCEEDED -&amp;gt; sales_amount + price * quantity&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Rankings can now be rebuilt from RDB metrics even if Redis is empty.&lt;/p&gt;&#10;&lt;p&gt;Weekly rankings can be computed by summing daily metrics over the required period, and monthly rankings work the same way. Neither requires replaying every original event.&lt;/p&gt;&#10;&lt;h2 id="score-calculation"&gt;&lt;a href="#score-calculation" class="header-anchor"&gt;&lt;/a&gt;Score calculation&#10;&lt;/h2&gt;&lt;p&gt;A separate batch or scheduler calculates scores. For example, it can read recent metrics every five minutes and apply the current weights.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;score = carry&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; + view_count * currentViewWeight&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; + like_count * currentLikeWeight&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; + ln(1 + sales_amount) * currentSalesWeight&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It then loads only the Top N products into a Redis ZSET.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ranking:top:daily:20260717&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Redis contains the results the API needs rather than the entire catalog. It serves as the fast serving layer, while the source data remains elsewhere.&lt;/p&gt;&#10;&lt;h2 id="what-this-design-improves"&gt;&lt;a href="#what-this-design-improves" class="header-anchor"&gt;&lt;/a&gt;What this design improves&#10;&lt;/h2&gt;&lt;p&gt;The biggest improvement is recoverability. If Redis data is lost, the Top N can be rebuilt from RDB metrics. If incorrect weights are deployed, the same metrics can be used to calculate corrected scores.&lt;/p&gt;&#10;&lt;p&gt;The second improvement is support for different ranking periods. With daily metrics, weekly, monthly, and yearly rankings become period aggregation problems. At larger scale, an RDB alone may not be enough; an OLAP store or separate aggregate tables could be needed. Retaining the source metrics leaves those options open.&lt;/p&gt;&#10;&lt;p&gt;Third, Redis memory usage becomes easier to control. Only the Top N needed for queries lives there. Full product metrics remain in the RDB.&lt;/p&gt;&#10;&lt;h2 id="what-it-costs"&gt;&lt;a href="#what-it-costs" class="header-anchor"&gt;&lt;/a&gt;What it costs&#10;&lt;/h2&gt;&lt;p&gt;This design isn&amp;rsquo;t automatically better in every respect.&lt;/p&gt;&#10;&lt;p&gt;Rankings become less immediate. A five-minute calculation schedule can introduce roughly five minutes of update delay.&lt;/p&gt;&#10;&lt;p&gt;The system also becomes more involved. It needs metric tables, upsert logic, score calculation batches, batch failure recovery, and Redis loading logic. There is more to manage than a direct &lt;code&gt;ZINCRBY&lt;/code&gt; on a ZSET.&lt;/p&gt;&#10;&lt;p&gt;RDB write load matters, too. Every user action can lead to a metric upsert. At higher traffic volumes, buffering, batch inserts, time-bucket aggregation, Kafka Streams, or an OLAP store may become worth considering.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="comparing-the-designs"&gt;&lt;a href="#comparing-the-designs" class="header-anchor"&gt;&lt;/a&gt;Comparing the designs&#10;&lt;/h2&gt;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Perspective&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Real-time Redis rankings&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;RDB Metric SOT + Redis Top N&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Priority&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Immediate updates&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Recovery and recalculation&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Redis role&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Real-time ranking store&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Query read model&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Source metrics&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Stored in Redis keys&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Persisted in the RDB&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Weight changes&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Recalculate only periods still in Redis&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Recalculate from retained RDB metrics&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Multiple periods&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;More keys add complexity&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Extend through metric aggregation&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Memory usage&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Grows when all products are loaded&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Controlled by loading only Top N&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Implementation complexity&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Relatively simple&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Requires batch and recovery policies&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Update delay&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Low&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Depends on the batch interval&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;The first design is a good starting point for fast daily rankings. User actions are reflected almost immediately, and the API read path is simple.&lt;/p&gt;&#10;&lt;p&gt;As recovery, recalculation, and rankings across longer periods become operational requirements, the RDB Metric SOT design becomes more compelling.&lt;/p&gt;&#10;&lt;h2 id="how-should-we-choose"&gt;&lt;a href="#how-should-we-choose" class="header-anchor"&gt;&lt;/a&gt;How should we choose?&#10;&lt;/h2&gt;&lt;p&gt;This exercise made me realize that the first question shouldn&amp;rsquo;t be which database to use. More useful questions are:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;How quickly must this ranking reflect new activity?&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Can we tolerate losing the ranking data?&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Do historical rankings need to be recalculated?&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Could the weight policy change frequently?&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Do all products need to remain in the ranking dataset?&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Will rankings stay daily, or expand to weekly and monthly periods?&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If the answers prioritize freshness, a Redis-centered design is simple and fast. If recovery, auditing, recalculation, and broader time periods matter, the system needs to retain source metrics.&lt;/p&gt;&#10;&lt;h2 id="evolving-the-design"&gt;&lt;a href="#evolving-the-design" class="header-anchor"&gt;&lt;/a&gt;Evolving the design&#10;&lt;/h2&gt;&lt;p&gt;The real-time Redis design initially seemed sufficient. Daily rankings needed to respond quickly to user behavior, and Sorted Sets supported both Top N queries and individual product ranks well.&lt;/p&gt;&#10;&lt;p&gt;I placed several limits on that design:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Use date-specific keys.&lt;/li&gt;&#10;&lt;li&gt;Give keys a TTL.&lt;/li&gt;&#10;&lt;li&gt;Carry over only the Top 100 products.&lt;/li&gt;&#10;&lt;li&gt;Keep metric ZSETs separate so weights can be recalculated for retained daily data.&lt;/li&gt;&#10;&lt;li&gt;Track processed events to avoid duplicate updates.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;These choices make it a starting point that favors freshness and simplicity. Once historical recovery or weekly and monthly rankings enter the requirements, evolving toward RDB Metric SOT + Redis Top N makes more sense.&lt;/p&gt;&#10;&lt;h2 id="closing-thoughts"&gt;&lt;a href="#closing-thoughts" class="header-anchor"&gt;&lt;/a&gt;Closing thoughts&#10;&lt;/h2&gt;&lt;p&gt;Designing rankings kept bringing me back to questions beyond sorting: which actions count as strong signals, how long to retain them, whether scores must be recalculated when policy changes, and whether an empty ranking after Redis failure is acceptable.&lt;/p&gt;&#10;&lt;p&gt;I began by focusing on how quickly Redis Sorted Sets could provide rankings. As the design grew, deciding what to preserve as source data became more important than deciding what to put in Redis.&lt;/p&gt;&#10;&lt;p&gt;The first design is a starting point for real-time daily rankings. RDB Metric SOT + Redis Top N is a possible next step as operational requirements grow.&lt;/p&gt;&#10;&lt;p&gt;I don&amp;rsquo;t think a good design has to implement every future requirement from the start. But I do want to be able to explain the limits of today&amp;rsquo;s choice and how it could evolve when those limits matter.&lt;/p&gt;&#10;</description></item><item><title>Handling Order Surges: Designing a Redis Waiting Queue</title><link>https://0andwild.com/en/posts/260710_waiting_queue_system_design/</link><pubDate>Fri, 10 Jul 2026 16:59:39 +0900</pubDate><guid>https://0andwild.com/en/posts/260710_waiting_queue_system_design/</guid><description>&lt;img src="https://0andwild.com/" alt="Featured image of post Handling Order Surges: Designing a Redis Waiting Queue" /&gt;&lt;h2 id="tldr"&gt;&lt;a href="#tldr" class="header-anchor"&gt;&lt;/a&gt;TL;DR&#10;&lt;/h2&gt;&lt;p&gt;Instead of sending every request straight to the order API, I first placed users in a Redis Sorted Set.&lt;/p&gt;&#10;&lt;p&gt;A scheduler removes up to 50 users from the queue every second and issues entry tokens valid for five minutes. Users poll their position until they receive a token, and the order API accepts only requests with a valid &lt;code&gt;X-Entry-Token&lt;/code&gt;.&lt;/p&gt;&#10;&lt;p&gt;After an order commits successfully, a local event listener receives &lt;code&gt;OrderEvent.Created&lt;/code&gt; and deletes the token. If the order fails, the token remains so the user can try again.&lt;/p&gt;&#10;&lt;p&gt;This is an initial design for controlling admission rate, rather than a complete production waiting queue. Atomic queue entry, recovery during token issuance, global throughput across instances, and polling load still need more work.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="overview"&gt;&lt;a href="#overview" class="header-anchor"&gt;&lt;/a&gt;Overview&#10;&lt;/h2&gt;&lt;ul&gt;&#10;&lt;li&gt;Goals:&#10;&lt;ul&gt;&#10;&lt;li&gt;Keep a sudden burst of order requests from reaching the order database all at once.&lt;/li&gt;&#10;&lt;li&gt;Let users check their queue position and estimated wait time.&lt;/li&gt;&#10;&lt;li&gt;Issue entry tokens to a limited number of users who can then call the order API.&lt;/li&gt;&#10;&lt;li&gt;Remove entry tokens only after a successful order transaction.&lt;/li&gt;&#10;&lt;li&gt;Keep the waiting queue and order domains independent of Redis implementation details.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;Settings used in this implementation:&#10;&lt;ul&gt;&#10;&lt;li&gt;Scheduler fixed delay: 1 second&lt;/li&gt;&#10;&lt;li&gt;Scheduler batch size: 50 users&lt;/li&gt;&#10;&lt;li&gt;Entry token TTL: 300 seconds&lt;/li&gt;&#10;&lt;li&gt;Rank: zero-based, using the Redis &lt;code&gt;ZRANK&lt;/code&gt; result directly&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="why-put-a-queue-in-front-of-the-order-api"&gt;&lt;a href="#why-put-a-queue-in-front-of-the-order-api" class="header-anchor"&gt;&lt;/a&gt;Why put a queue in front of the order API?&#10;&lt;/h2&gt;&lt;p&gt;A modest rise in order traffic may be handled by scaling out application servers or tuning the database connection pool.&lt;/p&gt;&#10;&lt;p&gt;A sharp burst at the start of an event is different. Even with more application instances, requests converge on the same database to lock inventory and coupon rows and save orders. Accepting more requests at the application tier can make database connection demand and lock waits grow even faster.&lt;/p&gt;&#10;&lt;p&gt;The queue&amp;rsquo;s purpose is to admit requests at a rate the downstream system can handle, while other users wait outside that processing path.&lt;/p&gt;&#10;&lt;p&gt;I chose not to store order request payloads in the server-side queue. The queue contains only &lt;code&gt;memberId&lt;/code&gt;; once admitted, the user sends the original order request again.&lt;/p&gt;&#10;&lt;p&gt;This avoids managing the state and retry policy of queued order payloads on the server. The client instead needs to poll its position and call the order API after reaching &lt;code&gt;READY&lt;/code&gt;.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="overall-architecture"&gt;&lt;a href="#overall-architecture" class="header-anchor"&gt;&lt;/a&gt;Overall architecture&#10;&lt;/h2&gt;&lt;figure class="mx-auto"&gt;&lt;a href="waiting-queue-architecture.png"&gt;&lt;img src="https://0andwild.com/posts/260710_waiting_queue_system_design/waiting-queue-architecture.png"&#10;&#9;&#9;&#9;alt="Overall architecture of the Redis waiting queue system" width="1400"&gt;&lt;/a&gt;&#10;&lt;/figure&gt;&#10;&#10;&lt;p&gt;The flow has three parts:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Register users in a Redis ZSET and expose their current status.&lt;/li&gt;&#10;&lt;li&gt;Let the scheduler remove users from the front and issue entry tokens.&lt;/li&gt;&#10;&lt;li&gt;Validate tokens at the order API and delete them through a local event after the order commits.&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;The layers have distinct responsibilities:&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Layer&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Responsibility&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Interface&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Waiting queue and order APIs, plus the order completion event listener&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Application&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;User authentication, queue status assembly, token issuance, and order admission validation&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Domain / Port&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;The &lt;code&gt;WaitingQueuePosition&lt;/code&gt; state model and Repository interfaces that abstract Redis&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Infrastructure&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Repository port implementations using Spring Data Redis&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Redis&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;A ZSET for waiting order and a String entry token for each admitted user&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="1-representing-queue-order-with-a-redis-sorted-set"&gt;&lt;a href="#1-representing-queue-order-with-a-redis-sorted-set" class="header-anchor"&gt;&lt;/a&gt;1. Representing queue order with a Redis Sorted Set&#10;&lt;/h2&gt;&lt;p&gt;The queue needs to answer three main questions:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Is the user in the queue?&lt;/li&gt;&#10;&lt;li&gt;What is their current position?&lt;/li&gt;&#10;&lt;li&gt;How many users are waiting in total?&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;A Redis Sorted Set can represent all three.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;key : queue:waiting&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;member : memberId&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;score : enteredAt&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Because &lt;code&gt;memberId&lt;/code&gt; is unique within the ZSET, adding the same user repeatedly doesn&amp;rsquo;t create duplicate members. Using &lt;code&gt;enteredAt&lt;/code&gt; as the score places earlier users ahead of later ones. &lt;code&gt;ZRANK&lt;/code&gt; returns their position, and &lt;code&gt;ZPOPMIN&lt;/code&gt; removes users from the front.&lt;/p&gt;&#10;&lt;p&gt;I use the &lt;code&gt;ZRANK&lt;/code&gt; value directly, so the first user&amp;rsquo;s rank is 0. Keeping both the internal model and API response zero-based avoids &lt;code&gt;+1&lt;/code&gt; and &lt;code&gt;-1&lt;/code&gt; conversions at every boundary.&lt;/p&gt;&#10;&lt;h3 id="preserve-the-position-on-repeated-entry"&gt;&lt;a href="#preserve-the-position-on-repeated-entry" class="header-anchor"&gt;&lt;/a&gt;Preserve the position on repeated entry&#10;&lt;/h3&gt;&lt;p&gt;A waiting user may refresh the page or call the entry API again. Overwriting the score with the current time would send them to the back.&lt;/p&gt;&#10;&lt;p&gt;The current implementation checks for an existing score and adds only users who aren&amp;rsquo;t already registered.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-kotlin" data-lang="kotlin"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;override&lt;/span&gt; &lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;enterIfAbsent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memberId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Long&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Double&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;val&lt;/span&gt; &lt;span class="py"&gt;member&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;memberId&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;redisTemplate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;opsForZSet&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;WAITING_QUEUE_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;member&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;redisTemplate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;opsForZSet&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="n"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;WAITING_QUEUE_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;member&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;}&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This preserves the score and rank across sequential repeated calls.&lt;/p&gt;&#10;&lt;p&gt;There is an important limitation: &lt;code&gt;ZSCORE&lt;/code&gt; and &lt;code&gt;ZADD&lt;/code&gt; are separate commands. Two simultaneous first-entry requests for the same user can both see no score and then write different scores. Strict idempotency requires a single atomic Redis operation such as &lt;code&gt;ZADD NX&lt;/code&gt; or &lt;code&gt;addIfAbsent&lt;/code&gt;.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="2-one-response-model-for-waiting-status"&gt;&lt;a href="#2-one-response-model-for-waiting-status" class="header-anchor"&gt;&lt;/a&gt;2. One response model for waiting status&#10;&lt;/h2&gt;&lt;p&gt;After entering, clients poll the same position API. It returns one of three states:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;code&gt;WAITING&lt;/code&gt;: the user is in the ZSET and has no entry token yet.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;READY&lt;/code&gt;: the user has an entry token and can call the order API.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;NOT_ENTERED&lt;/code&gt;: the user is neither in the queue nor holding a valid entry token.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Status checks look for the entry token first. The scheduler removes users from the ZSET before issuing tokens, so a &lt;code&gt;READY&lt;/code&gt; user is no longer in the ZSET. Checking rank first could incorrectly classify an admitted user as &lt;code&gt;NOT_ENTERED&lt;/code&gt;.&lt;/p&gt;&#10;&lt;div class="mermaid-box" style="max-width: 720px; margin: 1.5rem auto; overflow-x: auto;"&gt;&#10; &lt;pre class="mermaid" style="visibility:hidden"&gt;&#10;stateDiagram-v2&#10; [*] --&gt; NOT_ENTERED&#10; NOT_ENTERED --&gt; WAITING: enter() / register in queue&#10; WAITING --&gt; WAITING: position polling&#10; WAITING --&gt; READY: scheduler / issue token&#10; READY --&gt; READY: order rollback / retain token&#10; READY --&gt; NOT_ENTERED: order commit / delete token&#10; READY --&gt; NOT_ENTERED: TTL expires&#10;&lt;/pre&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;&lt;code&gt;WaitingQueuePosition&lt;/code&gt; uses the retrieved rank and total waiting count to calculate these response values:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;code&gt;status&lt;/code&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;rank&lt;/code&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;currentTotalWaitingCount&lt;/code&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;estimatedWaitSeconds&lt;/code&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;pollingIntervalSeconds&lt;/code&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;entryToken&lt;/code&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The estimated wait is the current rank divided by an assumed throughput of 50 users per second. The polling interval is one second with fewer than 100 users ahead, three seconds with fewer than 1,000, and five seconds otherwise.&lt;/p&gt;&#10;&lt;p&gt;Users near admission receive faster updates, while users further back make fewer Redis queries. Since the throughput assumption is fixed, this is only a rough estimate; it doesn&amp;rsquo;t reflect actual order processing speed.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="3-controlling-admission-rate-with-a-scheduler"&gt;&lt;a href="#3-controlling-admission-rate-with-a-scheduler" class="header-anchor"&gt;&lt;/a&gt;3. Controlling admission rate with a scheduler&#10;&lt;/h2&gt;&lt;p&gt;Sending all queued users to the order API at once would defeat the purpose of the queue.&lt;/p&gt;&#10;&lt;p&gt;&lt;code&gt;WaitingQueueScheduler&lt;/code&gt; calls &lt;code&gt;issueNextEntries()&lt;/code&gt; every second. The service uses &lt;code&gt;ZPOPMIN&lt;/code&gt; to remove up to 50 users with the lowest scores and issues a token to each.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-kotlin" data-lang="kotlin"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;issueNextEntries&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batchSize&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Long&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;waitingQueueRepository&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;popNext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batchSize&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mapNotNull&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;memberId&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;val&lt;/span&gt; &lt;span class="py"&gt;token&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;generateToken&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;entryTokenRepository&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;issue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memberId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;entryTokenTtl&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;entryTokenRepository&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memberId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;}&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Each token is a URL-safe Base64 encoding of 32 bytes generated by &lt;code&gt;SecureRandom&lt;/code&gt;.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;key : queue:entry-token:{memberId}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;value : generated token&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;TTL : 300 seconds&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The application doesn&amp;rsquo;t return the generated token directly. It stores it in Redis and reads it back, so a token that wasn&amp;rsquo;t stored isn&amp;rsquo;t treated as a valid admission credential.&lt;/p&gt;&#10;&lt;p&gt;The TTL serves two purposes. It prevents a token from granting indefinite access to the order API, and it eventually cleans up tokens even if deletion after order completion fails.&lt;/p&gt;&#10;&lt;p&gt;A very short TTL could expire while a user is preparing an order. A very long TTL leaves users eligible to enter long after they have stopped trying. The current 300-second value is the assignment&amp;rsquo;s default policy; a production value should be based on measured user behavior and target throughput.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="4-requiring-an-entry-token-at-the-order-api"&gt;&lt;a href="#4-requiring-an-entry-token-at-the-order-api" class="header-anchor"&gt;&lt;/a&gt;4. Requiring an entry token at the order API&#10;&lt;/h2&gt;&lt;p&gt;Order requests include an &lt;code&gt;X-Entry-Token&lt;/code&gt; header.&lt;/p&gt;&#10;&lt;p&gt;After authentication, &lt;code&gt;OrderFacade.placeOrder()&lt;/code&gt; compares the request token with the one stored in Redis. A missing, mismatched, or expired token results in &lt;code&gt;401 Unauthorized&lt;/code&gt; before order processing begins.&lt;/p&gt;&#10;&lt;p&gt;I put this validation in the application layer because eligibility to place an order is a precondition of the use case, beyond simply validating an HTTP header&amp;rsquo;s format.&lt;/p&gt;&#10;&lt;p&gt;This is a fail-closed design: when Redis doesn&amp;rsquo;t respond, orders are blocked. I chose to stop accepting orders instead of letting traffic bypass the queue and overwhelm downstream systems. That protects the server, but makes a Redis outage an order outage as well.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="5-deleting-tokens-outside-the-order-transaction"&gt;&lt;a href="#5-deleting-tokens-outside-the-order-transaction" class="header-anchor"&gt;&lt;/a&gt;5. Deleting tokens outside the order transaction&#10;&lt;/h2&gt;&lt;p&gt;I initially considered deleting the token directly inside the order creation method.&lt;/p&gt;&#10;&lt;p&gt;That would mix an external Redis operation into the database transaction flow. Deleting the token before commit is also problematic: if the order rolls back, the user loses admission even though the order failed.&lt;/p&gt;&#10;&lt;p&gt;The order flow already publishes &lt;code&gt;OrderEvent.Created&lt;/code&gt; on success, so I moved token deletion into a listener that receives this local event at &lt;code&gt;AFTER_COMMIT&lt;/code&gt;.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-kotlin" data-lang="kotlin"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nd"&gt;@TransactionalEventListener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;phase&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;TransactionPhase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AFTER_COMMIT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;handle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;OrderEvent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Created&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;waitingQueueService&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deleteEntryToken&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;memberId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This gives us two useful properties:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;A rolled-back order leaves the token available for another attempt.&lt;/li&gt;&#10;&lt;li&gt;A token cleanup failure cannot roll back an already committed order.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;A local event is not a durable message, however. If the process exits just after commit, or Redis deletion fails, the listener has no automatic recovery. The TTL provides eventual cleanup for now. Immediate cleanup would require a retryable event record or a separate cleanup job.&lt;/p&gt;&#10;&lt;p&gt;&lt;code&gt;AFTER_COMMIT&lt;/code&gt; also does not automatically mean asynchronous execution or exception isolation. The current listener runs in the same call flow as the order request. If a Redis deletion exception propagates, the HTTP response can fail even though the order has committed. Production handling needs an explicit policy for logging and alerting on those exceptions, alongside retryable cleanup work.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="technical-decisions"&gt;&lt;a href="#technical-decisions" class="header-anchor"&gt;&lt;/a&gt;Technical decisions&#10;&lt;/h2&gt;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Design area&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Choice&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Reason&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Waiting order&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Redis Sorted Set&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Unique members, score ordering, rank, and total count in one structure&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;ZSET member&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;memberId&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;One waiting entry per user&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;ZSET score&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Entry timestamp in milliseconds&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Give earlier arrivals priority&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Admission control&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Fixed-delay scheduler + batch size&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Limit the rate at which users reach downstream processing&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Admission credential&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Per-user Redis String token + TTL&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Separate waiting from readiness and expire old credentials automatically&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Status updates&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Client polling&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Expose positions through a simple HTTP API without a separate push channel&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Order validation&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Start of &lt;code&gt;OrderFacade&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Check admission before expensive order processing&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Token cleanup&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;AFTER_COMMIT&lt;/code&gt; listener for &lt;code&gt;OrderEvent.Created&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Clean up only successful orders and separate deletion from order rollback&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="trade-offs"&gt;&lt;a href="#trade-offs" class="header-anchor"&gt;&lt;/a&gt;Trade-offs&#10;&lt;/h2&gt;&lt;h3 id="admission-rate-is-controlled-but-concurrent-orders-are-not-strictly-bounded"&gt;&lt;a href="#admission-rate-is-controlled-but-concurrent-orders-are-not-strictly-bounded" class="header-anchor"&gt;&lt;/a&gt;Admission rate is controlled, but concurrent orders are not strictly bounded&#10;&lt;/h3&gt;&lt;p&gt;Issuing 50 tokens per second doesn&amp;rsquo;t guarantee that no more than 50 orders run at once.&lt;/p&gt;&#10;&lt;p&gt;Tokens remain valid for 300 seconds. Users admitted across several seconds can wait and then submit orders at the same moment. This design controls the token issuance rate, not actual in-flight order concurrency.&lt;/p&gt;&#10;&lt;p&gt;A strict downstream concurrency limit would need active slots, token claims, a semaphore, or another control. Another option is to admit users based on actual order completion rate.&lt;/p&gt;&#10;&lt;h3 id="polling-is-simple-but-creates-read-traffic"&gt;&lt;a href="#polling-is-simple-but-creates-read-traffic" class="header-anchor"&gt;&lt;/a&gt;Polling is simple, but creates read traffic&#10;&lt;/h3&gt;&lt;p&gt;HTTP polling is easy to implement on both sides and doesn&amp;rsquo;t require long-lived connections. But more waiting users mean more repeated &lt;code&gt;GET token&lt;/code&gt;, &lt;code&gt;ZRANK&lt;/code&gt;, and &lt;code&gt;ZCARD&lt;/code&gt; calls. Varying the polling interval by rank reduces that load.&lt;/p&gt;&#10;&lt;p&gt;At larger scale, I would consider client-side jitter to spread requests out, longer polling intervals, Redis read replicas rather than CDN caching for these reads, or SSE-based push.&lt;/p&gt;&#10;&lt;h3 id="separate-cleanup-does-not-guarantee-immediate-cleanup"&gt;&lt;a href="#separate-cleanup-does-not-guarantee-immediate-cleanup" class="header-anchor"&gt;&lt;/a&gt;Separate cleanup does not guarantee immediate cleanup&#10;&lt;/h3&gt;&lt;p&gt;An &lt;code&gt;AFTER_COMMIT&lt;/code&gt; listener reduces the responsibilities inside the order transaction. Its failures happen after the order succeeds, though, so the order cannot be rolled back. A local event also leaves no retry record.&lt;/p&gt;&#10;&lt;p&gt;The TTL eventually resolves the leftover-token state, but the token can remain valid until it expires.&lt;/p&gt;&#10;&lt;h3 id="a-redis-outage-stops-orders"&gt;&lt;a href="#a-redis-outage-stops-orders" class="header-anchor"&gt;&lt;/a&gt;A Redis outage stops orders&#10;&lt;/h3&gt;&lt;p&gt;Fail-closed behavior protects the database from requests bypassing the queue, while making Redis a required dependency of the entire order flow. Production design needs replication, Sentinel or Cluster, persistence, and failover policies as well.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="remaining-limitations-and-work-beyond-the-assignment"&gt;&lt;a href="#remaining-limitations-and-work-beyond-the-assignment" class="header-anchor"&gt;&lt;/a&gt;Remaining limitations and work beyond the assignment&#10;&lt;/h2&gt;&lt;h3 id="1-concurrent-entry-by-the-same-user-must-be-atomic"&gt;&lt;a href="#1-concurrent-entry-by-the-same-user-must-be-atomic" class="header-anchor"&gt;&lt;/a&gt;1. Concurrent entry by the same user must be atomic&#10;&lt;/h3&gt;&lt;p&gt;&lt;code&gt;ZSCORE&lt;/code&gt; followed by &lt;code&gt;ZADD&lt;/code&gt; is a check-then-act sequence. I verified concurrent entry by eight distinct users and sequential repeated entry by the same user. That does not establish atomicity for simultaneous first-entry requests from one user.&lt;/p&gt;&#10;&lt;p&gt;This should use the single &lt;code&gt;ZADD NX&lt;/code&gt; command.&lt;/p&gt;&#10;&lt;h3 id="2-users-can-be-lost-between-zpopmin-and-token-storage"&gt;&lt;a href="#2-users-can-be-lost-between-zpopmin-and-token-storage" class="header-anchor"&gt;&lt;/a&gt;2. Users can be lost between ZPOPMIN and token storage&#10;&lt;/h3&gt;&lt;p&gt;The scheduler removes a user from the ZSET before saving the token. If the process exits or the Redis write fails between those operations, the user has neither a queue entry nor a token and becomes &lt;code&gt;NOT_ENTERED&lt;/code&gt;.&lt;/p&gt;&#10;&lt;p&gt;Reading the token back after storage cannot recover a user who has already been popped.&lt;/p&gt;&#10;&lt;p&gt;A production design could atomically move users into a processing ZSET, then acknowledge them after token issuance succeeds. A recovery job could return users without an acknowledgment to the waiting ZSET after a timeout.&lt;/p&gt;&#10;&lt;h3 id="3-batch-size-is-not-a-global-limit-across-application-instances"&gt;&lt;a href="#3-batch-size-is-not-a-global-limit-across-application-instances" class="header-anchor"&gt;&lt;/a&gt;3. Batch size is not a global limit across application instances&#10;&lt;/h3&gt;&lt;p&gt;If every instance runs the scheduler, each can pop 50 users per second. &lt;code&gt;ZPOPMIN&lt;/code&gt; prevents duplicate removal of the same user, but four instances can issue up to 200 tokens per second in total.&lt;/p&gt;&#10;&lt;p&gt;Using batch size as a global throughput limit requires scheduler leader election, a distributed lock, a dedicated worker, or a Redis-based global rate limiter.&lt;/p&gt;&#10;&lt;h3 id="4-the-token-grants-temporary-admission-not-exactly-one-order"&gt;&lt;a href="#4-the-token-grants-temporary-admission-not-exactly-one-order" class="header-anchor"&gt;&lt;/a&gt;4. The token grants temporary admission, not exactly one order&#10;&lt;/h3&gt;&lt;p&gt;Token deletion happens after order commit. Two concurrent requests from the same user with the same token can both pass validation.&lt;/p&gt;&#10;&lt;p&gt;If one token must authorize only one order, validation and claiming it must be atomic. A simple &lt;code&gt;GETDEL&lt;/code&gt; at order start also removes the token when the order later rolls back. A more complete design needs states such as &lt;code&gt;READY → CLAIMED → CONSUMED&lt;/code&gt;, plus a release policy on rollback. Order-level idempotency keys deserve separate consideration, too.&lt;/p&gt;&#10;&lt;h3 id="5-estimated-wait-time-doesnt-reflect-real-throughput"&gt;&lt;a href="#5-estimated-wait-time-doesnt-reflect-real-throughput" class="header-anchor"&gt;&lt;/a&gt;5. Estimated wait time doesn&amp;rsquo;t reflect real throughput&#10;&lt;/h3&gt;&lt;p&gt;The current formula is &lt;code&gt;rank / 50&lt;/code&gt;. It matches the scheduler setting but ignores database latency, order success rate, users who receive tokens without ordering, and outages.&lt;/p&gt;&#10;&lt;p&gt;A moving average of recent token issuance and order completion rates, together with batch-size adjustments based on operational metrics, would provide a more realistic estimate.&lt;/p&gt;&#10;&lt;h3 id="6-arrivals-within-the-same-millisecond-are-not-strictly-fifo"&gt;&lt;a href="#6-arrivals-within-the-same-millisecond-are-not-strictly-fifo" class="header-anchor"&gt;&lt;/a&gt;6. Arrivals within the same millisecond are not strictly FIFO&#10;&lt;/h3&gt;&lt;p&gt;Scores use application timestamps in milliseconds. Multiple users can arrive in the same millisecond, and clock differences across application servers can change score order relative to actual arrival order.&lt;/p&gt;&#10;&lt;p&gt;For equal scores, Redis ZSET orders members lexicographically, which can also differ from arrival order. Strict FIFO would require a design using Redis &lt;code&gt;TIME&lt;/code&gt; with a sequence, or a separate incrementing value in the score.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="what-i-verified"&gt;&lt;a href="#what-i-verified" class="header-anchor"&gt;&lt;/a&gt;What I verified&#10;&lt;/h2&gt;&lt;p&gt;Using Testcontainers Redis and API E2E tests, I checked:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;First entry returns rank 0 and the total waiting count.&lt;/li&gt;&#10;&lt;li&gt;Sequential repeated entry preserves the user&amp;rsquo;s rank.&lt;/li&gt;&#10;&lt;li&gt;Concurrent entry by distinct users assigns unique ranks.&lt;/li&gt;&#10;&lt;li&gt;Responses represent &lt;code&gt;WAITING&lt;/code&gt;, &lt;code&gt;READY&lt;/code&gt;, and &lt;code&gt;NOT_ENTERED&lt;/code&gt;.&lt;/li&gt;&#10;&lt;li&gt;Tokens are unavailable after their TTL expires.&lt;/li&gt;&#10;&lt;li&gt;One scheduler run issues no more than the batch size, even with more users waiting.&lt;/li&gt;&#10;&lt;li&gt;Orders without an entry token are rejected.&lt;/li&gt;&#10;&lt;li&gt;A successful order commit deletes the entry token.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;These tests verify functional contracts. They don&amp;rsquo;t establish maximum capacity or the right batch size. To choose production settings, I would use a tool such as k6 to generate entry bursts and position polling together, while observing application latency, Redis command throughput, CPU, memory, and connection counts.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="alternatives-considered"&gt;&lt;a href="#alternatives-considered" class="header-anchor"&gt;&lt;/a&gt;Alternatives considered&#10;&lt;/h2&gt;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Option&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Benefits&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Costs&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Delete the token inside the order transaction&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;The full flow is visible in the order code&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Redis cleanup becomes part of the transaction&amp;rsquo;s responsibilities, and rollback handling becomes more complicated&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;strong&gt;Selected: delete through an &lt;code&gt;AFTER_COMMIT&lt;/code&gt; local event&lt;/strong&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Clean up successful orders and separate post-processing from the transaction&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Listener failures aren&amp;rsquo;t automatically recovered; tokens can remain until TTL expiry&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;The central decision was that token deletion should not determine whether an order succeeds.&lt;/p&gt;&#10;&lt;p&gt;Once an order has committed, preserving that success and cleaning up the token separately makes more sense than presenting a failure because cleanup failed.&lt;/p&gt;&#10;&lt;p&gt;The local event alone doesn&amp;rsquo;t provide that reliability, though. It separates responsibilities, while failure recovery still relies on the TTL. A cleaner structure and a more reliable production system are separate outcomes.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="closing-thoughts"&gt;&lt;a href="#closing-thoughts" class="header-anchor"&gt;&lt;/a&gt;Closing thoughts&#10;&lt;/h2&gt;&lt;p&gt;At first, I thought adding users to a Redis ZSET and returning their rank would complete the queue implementation.&lt;/p&gt;&#10;&lt;p&gt;Most of the difficult questions turned out to concern when a user becomes eligible to order, when that eligibility ends, and which state they should return to after a failure.&lt;/p&gt;&#10;&lt;p&gt;In this design, the ZSET represents waiting and the TTL token represents readiness. The order API checks the token, and an event after order commit triggers cleanup. This made the boundary between queue responsibilities and the order transaction clearer.&lt;/p&gt;&#10;&lt;p&gt;It also exposed gaps: atomic registration, recovery after pop, global throughput across instances, one-time token use, and wait estimates based on actual throughput. The assignment didn&amp;rsquo;t force all of these issues into view, but production use would require addressing them.&lt;/p&gt;&#10;&lt;p&gt;A waiting queue is an admission control system. It needs to move users at a pace the system can handle and preserve their state transitions even when something fails.&lt;/p&gt;&#10;</description></item></channel></rss>