We helped write the sequel to "Cracking the Coding Interview". Read 9 chapters for free

System Design: Coding Practice Platform Backend

Watch someone solve the design a coding practice platform problem in an interview with a FAANG engineer and see the feedback their interviewer left them. Explore this problem and others in our library of interview replays.

Interview Summary

Problem type

Design a Coding Practice Platform

Interview question

Design a system similar to a well-known competitive programming and interview prep platform. The system must support browsing and filtering problems by difficulty and language, providing an in-browser IDE, submitting code for evaluation, and returning results asynchronously. Stretch goals include supporting large-scale competitions with up to 100,000 simultaneous participants and a real-time leaderboard. Core challenges include securely isolating code execution environments, handling async job processing at scale, and choosing appropriate storage strategies for structured and unstructured data.

Interview Feedback

Feedback about Aerodynamic Lemur (the interviewee)

Advance this person to the next round?
Thumbs upYes
How were their technical skills?
3/4
How was their problem solving ability?
4/4
What about their communication ability?
3/4
Strengths & what went well Excellent overall performance for someone who said this was their first system design mock. You immediately identified the async pattern — submitting code, getting a submission ID back, and polling for results via Kafka — without any prompting. That is the single most important architectural decision in this problem and most candidates miss it entirely. Your API design was well-structured with proper REST verbs and clear request/response contracts. You correctly identified the need for secure isolated execution environments (sandboxing with cgroups, resource limits) as a non-functional requirement, which is a critical security consideration that many candidates overlook. Your scale estimation was reasonable and you used the numbers to inform design choices. The competition extension with Redis sorted sets and CDC from Postgres was a strong addition — you picked the right data structure for leaderboard ranking and explained why. The Kubernetes pod orchestration design was well thought through with HPA for autoscaling and health checks for crash recovery. Components were nicely isolated — pods interact only through the queue, not directly with the database, which is a good security boundary. Areas for improvement Storage was the weakest area. You defaulted to putting everything in Postgres, including test cases and problem descriptions, which are unstructured data that can be large files. The pattern here is polyglot storage: structured relational data (users, problems, submissions with their relationships) goes in Postgres, while unstructured content (problem descriptions with images, test case files, starter code templates) goes in blob storage like S3 with URL references stored in Postgres. This wasn't something you considered until prompted. The fix is to add an entity design step before jumping into APIs and storage — list your entities (Problem, User, Submission, TestCase), identify their relationships, then ask yourself what kind of data each entity contains. That naturally leads you to the right storage choices. The other gap was queue failover mechanics. You described health checks and retries at a high level but couldn't articulate when exactly messages get removed from the queue, what happens to in-flight messages when a pod crashes, or the difference between consuming-then-processing versus locking-then-processing-then-dequeuing. This matters because your current design ties pod lifecycle to job lifecycle (one pod = one submission), which breaks down at larger scale when you need to consolidate pods for efficiency. Advice for future interviews Before drawing any architecture, spend 2-3 minutes on entity design. List your core entities, their relationships (one-to-many, many-to-many), and what kind of data each holds (structured vs unstructured, small vs large). This naturally drives both your API design (entities map to REST resources) and your storage choices (relational for structured joins, blob storage for files). For Canva specifically, storage design matters a lot since their product is fundamentally about managing creative assets — images, templates, designs — which are all unstructured blob data with structured metadata. Showing strong polyglot storage thinking will resonate with their interviewers. Read up on the differences between Kafka and SQS for job queues. Kafka excels at high-throughput event streaming where losing occasional messages is acceptable. SQS is better for task queues where you need visibility timeouts, message locking, and guaranteed processing — which is exactly what code execution needs. Understanding this distinction and being able to justify your choice is strong senior-to-staff signal. Practice explaining queue failover patterns: lock an item so no other consumer picks it up, process it, then delete it only after successful completion. If the lock expires (pod crashed), the item becomes visible again for retry. This is a fundamental distributed systems pattern worth knowing cold. Open notes This was a genuinely strong first mock — comfortably at senior level with some elements trending toward staff. You have good architectural instincts and your component isolation was one of the best I have seen on this problem. The two areas to level up (polyglot storage and queue failover mechanics) are very learnable and will make a meaningful difference for your Canva interviews. For your remaining prep time, I would prioritize storage design patterns and queue reliability patterns over learning new system design problems — depth on these two topics will serve you across any question they throw at you. I also offer private coaching sessions outside the platform with more tailored, ongoing prep if you want to do a few more mocks before Canva. Happy to share details if you connect with me after this session.

Feedback about Admiral Hex (the interviewer)

Would you want to work with this person?
Thumbs upYes
How excited would you be to work with them?
4/4
How good were the questions?
4/4
How helpful was your interviewer in guiding you to the solution(s)?
4/4
He is awesome I want to have 3 - 4 more lessons with him thanks

Interview Transcript

Aerodynamic Lemur: Hey, can you hear me?
Admiral Hex: Yes, I can hear you. It's all right, thank you.
Aerodynamic Lemur: All right, so hi, welcome to your practice interview. We have 1 hour for this interview and And the mock itself will take like 45 minutes, and then we'll go over feedback in the last 15 minutes or so. Let's start with introductions, and maybe you can go first.
Admiral Hex: Oh yes, I'm just a senior backend engineer, just preparing for the senior level interview. And yes, we'll have an SBA level interview, 45 minutes. So yes, that's for me.
Aerodynamic Lemur: [REDACTED], my name is [REDACTED]. All right, uh, sorry, which company are you interviewing for?
Admiral Hex: Uh, I'm prepping for [REDACTED].
Aerodynamic Lemur: Can you write the name of the company here, please? [REDACTED]. Okay, okay, sounds good. So Do you know, um, what kind of questions to ask? Like, do you want me to focus on certain types of questions, or should I just ask a generic, like, systems design question?
Admiral Hex: I would say, uh, I will get 3-4 lessons. I still wanna you also get the next lessons too, if I know you by name, so that, uh, you know already who I am. By the next lesson, you know how to put push me harder, this and that. But I think it's, uh, their content provider, like, uh, images, uh, content for images, like image AIs and everything else. So everything gonna be like image related. Uh, they might ask a little bit of towards LLD, but I don't think so.
Aerodynamic Lemur: Okay, all right, cool. So I can briefly introduce myself as well. I'm [REDACTED]. I have about 12 years experience as an ML engineer. I used to work for [REDACTED] in the ads ranking team, and currently I work for [REDACTED] also in the ads ranking team. And I've been doing these interviews for more than 2 or 3 years and helped out like lots of candidates get through the process at different customers. Thank you. [REDACTED], right? Yeah, that's right.
Admiral Hex: Thank you, sir. And I'll take you for the next question if you as possible so that they assign you to me and then we work next week together. Fingers crossed.
Aerodynamic Lemur: Yeah, yeah, it sounds good. Uh, you know, you can also connect with me after this interview and I can send you a link for more, uh, lessons.
Admiral Hex: Yep, I would love to— one person to like run me end to end and as hard as possible. That's the best thing because every new person will require introduction and it will take time before that.
Aerodynamic Lemur: Yeah, yeah, makes sense. Cool. So let's get started if you're ready. If you can open the whiteboard, it's on the bottom left corner.
Admiral Hex: Okay, I click the toggle whiteboard.
Aerodynamic Lemur: All right, so I'll give you a high-level problem and I want you to design the solution, but feel free to ask clarifying questions as required, right? Yes. So I know that for [REDACTED] you said you wanted to image, but I'm going to ask you a different problem. I still think this problem is relevant for the kinds of things [REDACTED] is building, right? So yeah, I want you to keep an open mind with the problem and engage with it, right? If I was an interviewer at [REDACTED], I would ask this even though it's not— related directly to what [REDACTED] does. All right, okay, okay, okay. So we'll be designing [REDACTED] today. You might have used this, uh, product. It's a, a website where you can view programming problems and submit solutions to them.
Admiral Hex: Um, honestly, I think I've done it just like 3 days ago on my preparations. Do you want to go harder than this, or you think it's fine?
Aerodynamic Lemur: Yeah, I, I would still like to stick with this problem because I think it has a lot of depth. Uh, okay, so you practice it on your own. Now you can practice it in a mock setting, right? Yes. Yeah, let's go.
Admiral Hex: Uh, all right. Oh, what did I do? So this code, uh I start with the functional requirements. Okay. I will start with the functional requirements first, if you allow me to ask a couple of questions.
Aerodynamic Lemur: Yes, go ahead.
Admiral Hex: So, um, when we say lead code, uh, first thing we talk about is list of problems for the, uh, candidates to resolve, right? We have lists of problems. With their own complexities. Is it correct assumption?
Aerodynamic Lemur: Yeah, there are problems defined on the platform.
Admiral Hex: Yeah, problems, uh, with 3 levels of difficulty: easy, medium, and hard. And 8 to 10 different languages. That's the first functional requirement. OK to move on to the next one?
Aerodynamic Lemur: Yeah?
Admiral Hex: Then we have candidates. So candidates should be able to review, search, filter problems. And then they should be able— UI should provide the IDE to resolve selected problem. So on the candidate side, they should be able to review search field, the list of problems, and the UI should provide to the candidate the IDE to resolve the selected problem. Is that the correct functional requirement?
Aerodynamic Lemur: Requirement?
Admiral Hex: Yeah? Okay. So I think I'll add one more, which is, uh, leave Code Jam alone. Is the, uh, oh yes, after UI can provide the ID to resolve the selected problem, then the candidate should be able to submit, uh, their results and, uh, be able to poll Oh, sorry. Able to poll for the outcome. So the other one which is very big, [REDACTED] famous for, is competitions. So our competitions should be around, let's say, usually up to 100,000 participants. And then there should be a leaderboard and something like that as a functional requirement. Is this a correct one?
Aerodynamic Lemur: Yeah, sounds good.
Admiral Hex: Okay, I think that's enough.
Aerodynamic Lemur: The competition we can do as a stretch goal, so you can do stuff, and if you have time, we can go into competition.
Admiral Hex: Okay, thank you. Oh, so stretch goal is— yeah, okay. So then I focus on non-functional requirements. So on the non-functional requirements, I would go with, uh, NFR. So here what we get is, uh, the search, um, so I think usually, so DAU, MAU, I think we got around, let's say, daily active users, I would say maximum up to 100K, and then monthly is up to million. I don't think there is more than that for [REDACTED]. Is it the correct assumption?
Aerodynamic Lemur: Yeah, okay.
Admiral Hex: So then if it's, uh, DAU, MAU, then, uh, I think here latency and versus consistency. I think here latency, we go with availability. Sorry. Let's go with like this. I think I'll go with here with availability more than consistency. And I would say here that once the results are submitted, up to 2 to 3 seconds to get the final score. So we won't get into the like milliseconds to get the final score. But in terms of search and filter problems and UI render of IDE, here I would say we go to the 200 to 300 up to 500 milliseconds. Is it the correct assumption in terms of the availability or consistency for result submission? And then search and filter problems should be very quick. Yeah, so I would say latency is very small. Okay, um, so because the competition is the stretch goal, that's why the leaderboard is for now— it's a stretch goal, it's a leaderboard. Uh, what else can we do here? Oh yes, another NFR is So every submission should run in the secure isolated VM box type of environment, meaning CPU, memory, processes, file systems, everything Protected, protected, or I'll say boxed. Uh, is that a correct one too for the NFR? Yep. All right, so that's in terms of NFR. Did I miss any, or you okay for non-functional requirements?
Aerodynamic Lemur: Uh, yeah, I think I'm okay for now.
Admiral Hex: Yeah. Okay, so I do a bit of the back of the napkin, uh, computations. Which is for max 100K daily. Let's say each user does 10 submissions, so then it's 10^5, 10^5, and then it's, uh, oh, let's say each of them does 100 submissions per day, so then it's 10 7. When daily, it's going to be divided by 10^5. We get 100 submissions per second. So if each submission, uh, time limited for up to 5 seconds, then same time we get maximum 500 pods running. And let's say we kind of max limited up to 1,000 pods per second running at the same time. That's in terms of the one back of the napkin computations. Are you okay with this one?
Aerodynamic Lemur: Yeah, sounds good.
Admiral Hex: Okay, so now when we focus on availability over consistency, here what we do is, uh, on the client side— I just put this side UI— we do client-side polling. Every 500 milliseconds, uh, we poll the backend for the result. That's pretty much it in terms of the back of the napkin computations. If you're happy with it, I move on to API definition.
Aerodynamic Lemur: Yeah, that sounds good.
Admiral Hex: Okay, so now based on these two, I just try to see if I can make it even smaller than that so that we have it handy. Okay, so for the APIs, what we have is APIs. Um, one would be, uh, problems. It would be GET, would be problems. It would be, um, search— oh, sorry, problems. And so we're just getting the whole list. So that would be type It can be, let's say, easy, medium, or hard. Then it can be language, let's say Java, C++, and Ruby, and etc. And then it can be page, page number equals, let's say something. So what it would return us is, uh, the JSON array where each element would be problem ID, problem description, uh, then problem complexity, and let's say some kind of ranking, uh, in tasks. Let's say some kind of ranking. Let's say the percentage of how it will successfully resolve and this and that. So then the next one would be, of course, gonna be 200, would be get problems. So when the person wanna focus on the specific problem, it's gonna be by the problem ID. So then it would return the whole problem itself, problem ID. Then it would return us the detailed description. Then it would return us, um, Java or whatever language, uh, stub. Start would be like startup code, and the rest UI would pick up to render it in the IDE. That's going to be the get by problem ID, and then we would have the POST. So then you have again problems, then we again have the problem ID, And then here we would have the solution, and then it would be the, the code of whatever is the result. So the response would be, uh, 201 submitted, and then there would be, uh, I would say another GET problems Um, yeah, okay, I forgot this one here. Submission ID, uh, something which would be generated by our system. Oh, sorry, it's a— it's, it's in the body, in the response body. My apologies. So we get 201, and then response body, we do return the submission ID and the status which is submitted. So that's for the, uh, POST. And then we will get— get problems, uh, you would do problem ID, then we do submissions, then we would do submission ID, and where we would get 200, and then it would give us submission ID or submitted, failed, time limit exceeded, and etc. I think that's going to be enough for the APIs. Um, any comments about it or suggestions?
Aerodynamic Lemur: Yeah, I think the POST thing, right? So the last two APIs, you put them in problems, but probably they are their own entity, like the submission.
Admiral Hex: Um, uh, yes, by the way, because we're going to have a submission service, essentially that's a very good idea. That's easy. Submissions, problem ID.
Aerodynamic Lemur: So problem ID is probably then, uh, part of the payload of the— right, you can put problem ID. Yeah, and the GET can also become simpler, right?
Admiral Hex: Yeah, that's so true. We don't need all of these here. Yeah.
Aerodynamic Lemur: All right. Good. Let's go ahead.
Admiral Hex: Yeah, so that's in terms of the APIs. So just for reference, we just add the problem ID in the submission response. So that's in terms of the APIs. And now I can go on to the diagram if you have— thank you so much for your comments. So here we would be on the UI. First, I will try to break down the GET, but in summary, because of our numbers is 1,000 ports and then our searches are up to 100 submissions per second. Okay, I didn't calculate the searches. So probably in terms of the searches, so I just do another calculation. It's maximum up to 1,000 searches per second, which means any database can handle it easily. So one Postgres DB is enough. All right, so I start on the client side. So we have the client, um, then, uh, we have it pulling to our standard API Gateway Load Balancer. This one's going to be very important because, uh, API Gateway Load Balancer. Um, yep. So here, uh, we would have the search service. Or problem service, or whatever, however you want to call it. The idea of the search service is to provide the responses back and from the client to the, our DB. And yeah, this is our PostgreSQL. PG database. Here we're gonna have a problems list, problems, and then all the fields. So primary key is problem ID, and then we can do the secondary key, uh, problem type. And let's say we can have a secondary, sorry, not, yes, secondary index. So we just make this index by problem type and just make an index by language so that when we search it just makes our queries much easier. That's how we search our problems and that's how we get the —problems. So this is my, uh, problems endpoint. So client is hitting the API Gateway load balancer. Uh, if they needed, they do the security checks or rate limiting accordingly. And then it goes to the search service. Then it provides from the Postgres database, so up to 1,000 or even 10,000 requests per second, which can handle very easily read or write. This one's for search specifically, is gonna be all the reads. And then it's gonna return all the JSON responses. Are we okay for this search service or you have any questions or comments? Oh, sounds good. Okay, now the most complex one is, will be the submissions endpoint. Which is a submissions endpoint. So we're gonna again hit our API Gateway and Load Balancer, then we're gonna have the submission service, and then— sorry, the submission service, the first thing it's gonna do is create a record in the Postgres DB. Um, then it's gonna, um, PGDB submission record creation. So the second thing is it's gonna do is second, it's gonna connect to Kafka queue and create a record. So the reason why I'm adding up the Kafka queue is— I'm sorry, I will make the creation of the pods where the tasks run asynchronous. And the submission service will just submit everything to Kafka, and the Kafka will be connected to the, our Kubernetes cluster. So here we'll have our Kubernetes cluster, and it will do manage all the virtual VMs Ops, sorry. Where is that? In a secure way and autoscaler, it will have the HPA. So, um, yeah, and it will create a record and then it's a Kafka queue. And then the reason why we're doing it, because the HPA, it will take time to, let's say, jump from 100 requests to 1,000 to scale in. That's why it will just need to do the asynchronously, and it will be connected to the Kafka, and then based on the log and request, It will create all the pods, and then the idea is that it will publish another Kafka, let's say, response with the job ID or pod ID, or pod ID, whichever. And then the service will be pulling back the Kafka for created job or pod ID. Once it's done, so here the second step will be this one, uh, second step Kafka, uh, produce and consume. It's going to be here. Then, uh, our submission service can, by that pod ID, let's say, call the Kubernetes API. That's going to be our step number 3: call Kubernetes API by pod ID or job ID, whichever that comes back, right? So Kubernetes, uh, Cluster definitions will be responsible for this, uh, forking bombs. Then there is, uh, security of the processes, uh, I think cgroups and others around, uh, limiting the pod. And after the, uh, Kubernetes responds with the job ID and job result Then the step number 4 will be returning the results to the UI. For return— no, sorry, my bad. We will update the database and number 4 will come here. And for Postgres, if the results the results of, uh, kubespot job results, whatever it is, uh, success, failure. Okay, another thing is that these ones will be all time limited so that they, uh, killed automatically if they go over time. Um, what the client will do is this one will be submissions POST. And then for the GET, every, uh, let's say 500 milliseconds, we will have the submission ID which will come right here. And through here, let's say it will come to the database. And then every 500 milliseconds, it will be able to return the submission status. So that's shortly my system design diagram. Um, yep. Do you have any comments or any suggestions for now?
Aerodynamic Lemur: So yeah, I'm just, uh, one thing is what exactly is the pod doing? Can you be more explicit about it? So what's the job it's running exactly?
Admiral Hex: Right. Okay, that's a very great comment. I think, uh, the job of the pods is— job of the pods is first scale up the OS, then scale up the environment, like language might be C. How does it know the language? Sorry?
Aerodynamic Lemur: How does it know the language?
Admiral Hex: Okay, so it comes up in the Kafka request. Kafka request would have everything.
Aerodynamic Lemur: So what happens in the Kafka request, like?
Admiral Hex: So our cluster would be pulling the Kafka, and then in the Kafka message, it would automatically obtain the submission code and then the language, and then it probably would have like Presets.
Aerodynamic Lemur: You're missing something here, right? Like, to actually run the code, you need the submission code, you need the problem ID, so you need the language. And then from the problem ID, what, what things are you pulling, right? Like, what's actually running? Like, you're running the code against what?
Admiral Hex: So I'm expecting that, uh, Postgres first call of the submission service would fetch everything around the problem and submission, and then when it creates Kafka produce message, it would, uh, submit these details here, and then that would provide all of these details to our Kubernetes cluster HPA. And then what I'm—
Aerodynamic Lemur: trigger the job actually talking about is test cases.
Admiral Hex: Yeah, yeah, yeah. So that's why, uh, okay, thank you for the comment. That's what I'm saying, we will have the preset, uh, ah, yes, we can have the preset image with all the cases. Yes, and I didn't mention it specifically, my apologies. So here the problem, you have the test cases too, uh, in this field. So when, uh, we pull everything about test cases from?
Aerodynamic Lemur: Like, where are the test cases stored?
Admiral Hex: The test cases are stored in a Postgres database, as, as I'm showing on the screen. And when we're doing the first call to the submission service and it calls Postgres, it picks up all the information about the problem, including test cases for that language.
Aerodynamic Lemur: All right, now just, uh, follow up on storage. So things like the test cases and the problem descriptions, do they— is it the best idea to store them in Postgres? That's the question, right? Like, is there a more appropriate place to store them?
Admiral Hex: So I would say it depends on the numbers. So I for the minimal levels where we are, we can calculate everything and usually Postgres databases easily can go to 100 terabytes, right? So if we approximately say, yeah, I think, I think so the small number of problems and small scale still okay, but yes, one thing you see, you hinting me and suggesting, which is also a great idea. What we can do is we have the S3 bucket, for example, and then all the blob, blob data, which is binary data related to the problems. We can create here, yeah, test cases. I think can be linked here by the problem ID. And then that would have the URL, which is also, uh, can be test cases S3 URL.
Aerodynamic Lemur: I think it's not so much about the size of the storage as it is about the format, right? So test cases are more like, you know, files in general. Problem is they're unstructured data, right? So, you know, it's just easier to work with it if it's an S3 Uh, because it doesn't do well with that kind of info, right? You'll have to force it into like big records, and then, you know, Postgres has issues with that. Like, yeah, Postgres is good storing like billions of records where each record is more or less small, but if you're storing like, you know, a few text files in Postgres That's an anti-pattern, right? Yeah. Okay, let's talk about— yeah, let's talk about the queue. So can you comment briefly about like all the failure recovery scenarios around the queue? Like what if a pod crashes, etc.? And like how does that Retry, failover, teamwork.
Admiral Hex: Right, I actually didn't focus on it here, and if you allow me, I just think a second in terms of retry and failover. Yes, the pods can crash, and let's say pod crashed, uh, so pods crash, right? Pods crash because— so they can go over time, right? So if they go over time, our step 3, which is called kubes-api, uh, will kind of set the status to time limit exceeded, and that's it. There's no need for retry. Now if it just pod crashes because of some reason like not enough memory or space or some other things. Uh, retry should happen through— that's a great question. Uh, probably again our step 3. I think our submission service, submission service. If pod crashes— okay, I see. So there are two things if pod crashes. Our Kubernetes cluster will do the health check, and if pod crashes, uh, will the, uh, restart the new replica. This is one thing. Now let's say they give it, let's say, 3 tries maximum, and then if all of these 3 fail when we are on our submission service, what we do is we do see that port has crashed. We probably create a new Kafka message and then maybe retry it 2, 3 times again. If again all of them fails, then we do the DLQ. DLQ will be the dead letter queue, which gonna be kinda UI to customer service. And then this is a very specific case scenario where the customers would have to analyze the dead letter queue about all the logs, the issues of the admins, and this and that. I hope I could answer the question.
Aerodynamic Lemur: Yeah, I mean, just— I think one thing is like, when are you actually removing stuff from the cubes?
Admiral Hex: Oh, from the cubes. So on the cubes, yeah, so on our images they have the, uh, can, uh, time limit to run. So they killed automatically, uh, batch number 1, or when they finish, they just cleaned up in the cubes. That's one, one way. In terms of cleanup.
Aerodynamic Lemur: I think what you have not addressed is the crashes, right? I mean, health check is fine. If pod crashes, restart, right? But to restart it, you need the payload in the Kafka queue, right? Right. Um, either is that stored in the DB, like, and as in, you know, the thing with Kafka queue is you like the point where it's fuzzy is, you know, okay. One thing you can do is every time you make a new pod, it consumes from the queue. Right. So the message is no longer. You can— yeah, now the problem is if it crashes, you'll have to put a new message in somehow. So you're saying I'll do a health check and then I'll look at my DB and take all the info from there and put a new message. Okay, that's fair enough. I think I would say another part—
Admiral Hex: I would say I think Kubernetes might see the, uh crashed pod definitions. But you're right, if the pod crashes and we don't have any definition, I don't think I put clearly what should we do, and I'm just thinking about it. I think you're right.
Aerodynamic Lemur: So yeah, no worries, I think that is fine. Um, yeah, and just quickly, let's talk about competitions and leaderboard.
Admiral Hex: Well, that's a great question too. So what competitions will do is, uh, they will mostly, uh, adapt our submissions, uh, maybe the competition ID in the payload. Let's see where it can be. Problem ID and then— competition ID and something. So the first thing it does is it adds up into the scale. Usually what happens is, let's say, the— wait a minute— 100K participants, it means we will have around 100K submissions, it might— we might have this much scale submissions per second, right? So we have to after autoscale our submission service, and then we might have to add some sharding to database. So I just set up some kind of scaling of submission service a little bit, so it's a little bit rude. And then here, we might need to do some sharding by, let's say, by— sharded by, let's say, user ID so that it's equally spread out in terms of the reads and writes. Let me read from— And we read from— Another one is we probably will have to do the Reddit Search Leaderboard. So the reason being because it's 100,000 people. This is a search service. —adapt Redis cache. So Redis provides the data structure which is called sorted set. Redis data, Redis cache, and it will have the leaderboard by user ID and sorted set. That's the structure. Yep. So every time we do the update to the DB, what we can do here is we can pick up from here and then do the CDC to our Redis leaderboard and let our Redis update Oops, sorry. Yep, sorry, it's my— a little bit like this. So we will have the CDC, Change Data Capture. Once the submission's results are up, it will publish to our Redis leaderboard, and then the UI can just with the submissions here, we will also get back the— I think this one, and you get the leaderboard. The reason why I'm using the Redis cache, because it has the data structure ready and it will handle the load of 100,000 records per second easily, and that structure is ready to calculate the leaderboard automatically. I hope I could answer the question, sir. Hello, hello, sorry, I think I lost you.
Aerodynamic Lemur: Yeah, yeah, sorry, I was on mute. Cool, sounds good. All right, we can talk about like the feedback now. Uh, how do you think you did?
Admiral Hex: Um, I think I definitely hit the middle level, plus minus. I'm not sure for the higher level. This is my first ever system design interview in my life. So yeah, now maybe you can suggest me how I did, but definitely probably not the junior, maybe better than middle.
Aerodynamic Lemur: Yeah, I think this was like a solid senior interview. Thank you. Um, so yeah, I think you did a great job, especially like, uh, like you went directly for the queue design with like the async polling pattern. That's great, right? And then, uh, yeah, you also noticed that we need isolation and security on the pods, which is something most people miss. Yep. Um, you were able to complete the leaderboard design at the end with sorted set, change data capture, pretty much very reasonable design. And also, like, I think you handled the Kubernetes pattern pretty well. The thing I like about your design is like It's all very isolated, right? It's only like API interaction, so the Kubernetes pods can't write to the DB. No, it's all been created. Um, so yeah, that's kind of good. The one thing I would— two things I would say actually. Okay, let's start with the most major thing, which is the storage, right? So you just jumped into Postgres storage, which is reasonable. But you need to think about Postgres plus S3, like polyglot storage design, right? Uh, different kinds of data require different storage requirements, so just take that into account. So storage is an important area, right? Which I would say was probably your weakest stage in this interview, right? Everything else was pretty much solid. Maybe storage You had to be prompted a bit.
Admiral Hex: Thank you, sir. I agree 100%.
Aerodynamic Lemur: Hey, can you hear me? Yes, you're back now.
Admiral Hex: Thank you, sir. Yeah, so Like I was saying, storage.
Aerodynamic Lemur: Again, keep saying that 2-3 times. That's the main thing, all right? Apart from that, everything's good. Now there are a couple of minor nits. Okay, one is like I had to get you the, the problem versus submission thing in the API. Oh, right. Then you didn't model the test cases. You didn't talk about test cases at all. I had to ask you. So typically before we do API or high level, one thing that people do is entity design, right? So say entities are like problems, submissions, and a problem might have test cases, right? And then there are users. So what that helps you is reason abstractly about like, okay, what are the things you're dealing with?. And then you can be like, oh, I need to do lots of joins between problems, submissions, users. So because of that, I'll use a relational DB, right? So you can motivate the choice. And you can be like, oh, there's like structured and unstructured data here. So there's problem descriptions, there's test cases, all of that is unstructured. So I can use different storage choices. So I can put all the structured data in a relational DB because I need to do joins between problem submissions users, test cases, etc. And you know, there's all tags or whatever, right? There are all kinds of things, right? So relational DB pattern works. And then you can say, but relational DB, I can't store like text data, images, all of that stuff. So problem descriptions might have images, text. I'll store all of them in S3. And test cases might be big files, so I could also store them in S3. And then when I'm submitting stuff, I know the problem ID, I have the submission ID, I know the test case IDs, so I can send them all of that information to Kafka, and then my pod can fetch that from S3. Or, you know, like, you can have a facilitator that pulls all of that stuff, packages it in the in the, uh, the payload to the Kubernetes, and then it runs, right? Um, so yeah, so I mean, actually, did you get the difference between the way you presented it? It's a small difference between the way you said it and the way I said it. The only difference is I have motivated my choices, right? I've said There are 3 entities, there are a lot of joins, so I'll use relational. There's unstructured data and structured data, I'll use structured data in, in Postgres, unstructured data in S3, right? So I've motivated the choices, right? Yes. The solution is roughly the same as what you said, right? And what helps— this helps you, then you're like, oh, submission is like an entity in itself, so I should have its own POST and GET, right? You know, Like just a bit of abstract reasoning allows you to make good choices on both storage, on APIs and all of that.
Admiral Hex: Thank you, sir. I really like it. I really like what you're saying. I really appreciate it. Thank you so much. All right.
Aerodynamic Lemur: Yeah, I think that is— and then queue, you know, is always tricky because, uh, like you can you can just say, oh, I'll put a queue in. I think you went into good amount of detail, but always remember with the queue pattern, like here again you just said Kafka, right? You can use SQS as well. Now what's the difference between them, right? The difference is more sort of persistent, right? So if you have this kind of use cases where you have work packages, and then, you know, you can have failures and you need retries and you need some more stronger consistency, let's say, probably SQS is a better choice, right, than Kafka, which is why I was pushing you on, you know, like, okay, like, how are you, uh, removing stuff from the queue and stuff, right? Because Kafka is good for analytics events where it's like, you know, you push it in one place and you get it from the other. But then if you lose it, like if you don't process, fine, it's just one analytics event, who cares, right? But it's probably stronger. Can, can you help me to—
Admiral Hex: let's say if you were me, how would you play it out? You would say I have to use the queue because the scale, uh, you have to be asynchronous.
Aerodynamic Lemur: Queue is just for air gapping in a way, right? Like basically your jobs are gonna take a few seconds to run. You don't want the user request to hang. Yeah, yeah, yeah. It's like, oh, just write it to queue and return async. It's an async pattern, right? Fine, that's okay. Now, what are you actually using the queue for? You put some stuff in there on the basis of which you're running a job in a pod, right? So wasting resources— like, not wasting, you're using resources for doing the work in the queue. So then you need to talk about failover right now. The main thing with failover is when do you remove stuff from the queue? Because if you remove stuff before your job is done Okay, then you need to put things back in the queue somehow, but then like that's the pattern you used where it's like something monitoring the job, but then all of this depends on having one pod per submission, right? So this only works because you have one job per submission, one pod per submission, which may not be a good use of resources. So tomorrow, if you say I have too many pods, I've like, okay, now you're scaled out to like millions of users. They're all doing like thousands or tens of thousands of millions of submissions. You have 10,000, 100,000 pods. Each pod is not doing anything very major. You're not reusing pods. So you're like, there's just too much overhead in the system. So you want to consolidate your pods, but then your failover doesn't work because how do you know the job status? You're relying on pod status, right? So this is like the part where you— there is like a problem in your design, right? You want to get the problem, right?
Admiral Hex: Yes. Why did you say that SQS would be better than Kafka? My— or you motivated something around SQS? To failover? I missed that part a little bit.
Aerodynamic Lemur: Yeah, it's just that, you know, Kafka, SQS has more guarantees and just the API is easier to work with. So what you can do in SQS is you can say, you know, there's an item, I'm working on it, so someone else shouldn't dequeue it, but don't remove it from the queue either. Right, because I don't know like whether I've completed this or not yet. You can say that kind of stuff with SQS. You're like, oh, this item is in the queue, I have taken it, but don't remove from the queue, keep it because once I'm done I'll tell you remove it, like that kind of thing, right?
Admiral Hex: Ah, okay, okay. So queue failover is a good topic for me to brush up.
Aerodynamic Lemur: Yeah. Yeah. And also like, I think whole pod ID equal to job ID thing, just think about like, what are the limitations of that? I didn't go too much into it, uh, for you because I mean, this is where it starts to move towards the higher end of senior or staff level. But you see, even with this simple design, which you think you almost got, Um, there's still, uh, there's still a limitation, right? So that I just want you to understand that, like, there is like a lot of depth with this queue pattern, right? So if you talk about like queue failovers, different types of queues, what is the difference between Kafka and SQS, like, you read up a bit more about that, you know, that will be helpful for you to think to brush up on, read more, understand more. There's a lot of things you can do there, right? Yeah, so you can have like one pattern is you have locks. So you take an item from the queue, lock it so no one else can dequeue it, but it's still in the queue. Then once you're done, then you dequeue it, right? Wow. So all of these come with trade-offs. Yeah, yeah, because you know, the thing is, if you dequeue stuff before it's done, then it's a problem.
Admiral Hex: That's a problem. Thank you, sir. Thank you so much.
Aerodynamic Lemur: I really appreciate it. Cool. Yeah. Um, for your interview with [REDACTED], don't worry too much. You are, you're doing well. You should be able to get through. Um, but yeah, like I think storage and queues, these are good topics to read more on that helps you improve as an engineer. Helps you do better in your future interviews as well. So treat it more like a stretch goal or something to improve upon, but not necessarily blocking your interview. Like, you're doing well overall, you should be able to pass the interview. So go in with confidence.
Admiral Hex: So, so when you give me your link, I will book you and then you run me through some other patterns like, uh, Instagram or ticket booking, you know, so that you touch all my other areas. Yeah, awesome, because you already know me, you already know where I can fail, what new things to ask. Thank you, sir.
Aerodynamic Lemur: All right, yeah, yeah, overall, like, you know, really good performance. So yeah, that's the final message I want to leave you with. You should be able to do well in the interview.
Admiral Hex: Yeah, I still want to do two. Thank you. I want to guarantee the result. Thank you, sir, and we'll do another couple. I'll look for your email. Thank you, sir.
Aerodynamic Lemur: All right, yeah, so actually at the end, uh, like when you're submitting your feed, you know, you, you can rate me as well, but there's an option over there to like, you can get yourself introduced to me. So if you click that, it will connect us and I'll have your email ID and then I can reach out to you over email to set something up, right?
Admiral Hex: Okay, I got it. I'll try to search for you. Yep. Because I need another one somewhere on Monday for sure, or Wednesday the latest. Yeah, a couple on the weekend again. Thank you, sir. So your name is—
Aerodynamic Lemur: we can totally do— once again, you, you— [REDACTED]. [REDACTED]. Yeah, you'll get my email if you check the box at the end, right? So yeah, we can, we can be in touch over email.
Admiral Hex: All right, thank you. Thank you.
Aerodynamic Lemur: My name is [REDACTED].
Admiral Hex: Oh yeah, nice. I put a note on it. Thank you, sir.
Aerodynamic Lemur: Thanks a lot. Thank you, sir. Bye-bye.

We know exactly what to do and say to get the company, title, and salary you want.

Interview prep and job hunting are chaos and pain. We can help. Really.