Testimonials
Services
System Design for the LLM Era
About me
- Architecturally Speaking | Sampriti Mitra | Substackhttps://architecturallyspeaking.substack.com/

Frequently asked questions
What is system design in software engineering?
System design in software engineering is the process of defining an application's architecture, components, data flow, and technology choices so it meets goals like scalability, reliability, low latency, and cost efficiency. It is the skill of answering questions like "How should this app serve millions of users?" — covering databases, caching, load balancing, queues, and APIs — rather than just writing individual functions.
How do I prepare for a system design interview?
Build the fundamentals first — load balancing, caching, SQL vs NoSQL, sharding, replication, and message queues — then practise designing common products like a URL shortener, chat app, or rate limiter out loud within 35–45 minutes. Walk into the system design interview with a repeatable framework: clarify requirements, estimate scale, sketch the high-level design, deep-dive into one component, and close with bottlenecks and trade-offs. Mock interviews with feedback accelerate this far more than passive reading.
What are the most common system design interview questions?
The most frequently asked system design interview questions include designing a URL shortener, a WhatsApp-style chat, an Instagram/Twitter feed, a rate limiter, a video streaming platform, a notification system, and a ride-hailing or food-delivery backend. In India, UPI-style payment system design is especially common at fintech companies. These questions are open-ended, so interviewers evaluate your trade-off reasoning rather than a single correct answer.
What is a good system design roadmap for beginners?
A practical system design roadmap looks like this: first solidify networking, database, and OS basics; then learn horizontal scaling, load balancing, caching, and CDNs; next go deep on data storage — SQL vs NoSQL, indexing, partitioning, replication; then study consistency, availability, and CAP trade-offs; and finally practise designing real systems end-to-end while reading engineering blogs. Plan for roughly 8–12 weeks of consistent effort alongside your regular DSA prep.
Is the System Design Primer enough for interview preparation?
The System Design Primer is one of the best free starting points because it consolidates core concepts, design templates, and case studies in one place. On its own it is rarely enough, though — combine it with at least one in-depth book and hands-on design practice, because the primer teaches concepts but not how you perform under real-time pressure in front of an interviewer.
Which system design books are worth reading?
The most recommended system design books are Alex Xu's "System Design Interview – Vol. 1 & 2" for interview-focused practice and "Designing Data-Intensive Applications" by Martin Kleppmann for deep fundamentals. If you are building AI-native products, "System Design for the LLM Era" by Sampriti Mitra is a useful recent addition, as it focuses specifically on integrating LLMs into production-grade systems — a gap most classic books do not cover yet.
Is a system design course worth it?
A system design course or paid mentorship is worth it when you need structure, accountability, or expert feedback on your design trade-offs — especially if interviews are only a few weeks away. If you are self-disciplined and have time, free resources plus a good book can cover most of the syllabus; the real differentiator is practising complete designs and getting them reviewed.
What are distributed systems in computer science?
Distributed systems in computer science are groups of independent computers that communicate and coordinate over a network to appear as one coherent system to users. Examples include distributed databases like Cassandra, streaming platforms like Kafka, and the backend behind any payment app processing millions of transactions. Their defining challenges are partial failures, consistency, fault tolerance, and latency.
How to learn distributed systems from scratch?
If you are figuring out how to learn distributed systems, start with prerequisites — computer networks, operating systems, and the CAP theorem — then work through a rigorous distributed systems book such as "Designing Data-Intensive Applications" or "Distributed Systems: Principles and Paradigms". Follow that with a structured course or consistent reading, and finally build small projects, because these concepts only truly click when you build and break real systems.
What are some good distributed systems projects for beginners?
Beginner-friendly distributed systems projects include a replicated key-value store, a simple message queue, a distributed rate limiter, a sharded URL shortener, and a leader-election service. The fastest way to learn how to build distributed systems is to deliberately break your own project — kill nodes, inject network latency, and observe how replication and consensus algorithms recover.
Distributed systems vs microservices: what's the difference?
When people frame distributed systems vs microservices, they are comparing a broad discipline with a specific style: distributed systems is any setup where multiple machines cooperate over a network, while microservices is an architectural pattern that splits one application into small, independently deployable services. Every microservices application is a distributed system, but not every distributed system uses microservices.
What are the most common distributed systems interview questions?
Most distributed systems interview questions revolve around the CAP theorem, consistency models, replication vs partitioning, leader election, and consensus algorithms like Raft and Paxos. Expect scenario probes as well — what happens during a network partition, how to achieve exactly-once processing, how to design a distributed message queue, and how to handle clock skew across data centres.
Can you explain how LLM architecture works in simple terms?
Here's how LLM architecture works in simple terms: your text is split into tokens, tokens are converted into embeddings, and stacks of transformer blocks use self-attention so every token can weigh the context of every other token; an output layer then predicts the next token one step at a time. This same design explains most behaviours you deal with in production — context-window limits, latency, cost per token, and why retrieval (RAG) is wrapped around the model.
What are the different LLM architecture types?
The main LLM architecture types are encoder-only models such as BERT (understanding, embeddings, classification), decoder-only models like the GPT family (text generation — the basis of most modern chat LLMs), and encoder-decoder models like T5 (translation and summarisation). Newer variants include mixture-of-experts designs for cheaper scaling and multimodal architectures that process text, images, and audio together.
How to structure an LLM prompt to get reliable outputs?
If you are unsure how to structure an LLM prompt, use this order: a clear role, relevant context, the specific task, constraints, the desired output format, and two or three examples. For production applications, also keep system instructions separate from user input, keep each prompt tightly scoped, and validate the model's output before your system acts on it.