{"$note":"Generated by scripts/build-manifest.js — do not edit. Run `npm run manifest`.","site":{"name":"Learn AI Bytes","description":"Interview-prep notes on AI models and LLMs — what they are, their types, and real-world examples. Search a topic or explore the full guide.","url":"https://learnaibytes.com/","author":"Kallol Chakraborty","authorUrl":"https://www.linkedin.com/in/kallol-chakraborty-9728a699/"},"phases":[{"id":"llms","title":"AI & LLMs","description":"A sequential AI → ML → Discriminative/Generative Algorithms → Deep Learning foundation, then how LLMs work: inference, prompting, forward and backward propagation, transformers, and caching.","order":1,"guides":[{"id":"what-are-llms","title":"AI, Machine Learning & Deep Learning","description":"A sequential learning path: AI fundamentals → Machine Learning (supervised/unsupervised/RL) → Discriminative vs Generative models → key algorithms for each (LogReg, SVM, Trees, kNN, Naive Bayes, GMM, VAE, GAN, Diffusion, Transformers) → Deep Learning architectures (MLP, CNN, RNN, Transformer).","icon":"auto_awesome","order":1,"sections":[{"id":"what-is-ai","title":"What is Artificial Intelligence (AI)?","icon":"smart_toy","order":1,"file":"okf/llms/what-are-llms/what-is-ai.md","body":"---\ntype: Section\ntitle: What is Artificial Intelligence (AI)?\ndescription: AI, Machine Learning & Deep Learning - What is Artificial Intelligence (AI)?\ntags: [what-is-ai,what-are-llms,llms]\ntimestamp: 2026-09-05T22:05:00.000Z\nsection: what-is-ai\nguide: what-are-llms\nphase: llms\nicon: smart_toy\norder: 1\n---\n\n# What is Artificial Intelligence (AI)?\n\n**Icon:** smart_toy\n\nAt its core, **Artificial Intelligence (AI)** is the science of making machines do things that would require intelligence if done by humans. It is the broad **umbrella term** for any system that can perceive its environment, reason about it, and take action to achieve a specific goal.\n\nThink of AI as the \"Smart Assistant\" in your phone or the recommendation engine on a streaming service. It's not a single algorithm; it's the entire field of study aimed at building intelligent agents.\n\n### The Evolution of AI\nAI isn't new—it has evolved drastically over the decades, moving from rigid rules to systems that learn organically:\n\n| Era | Approach | Example |\n|---|---|---|\n| **Symbolic AI (1980s)** | **Rules & Logic:** Engineers hand-write strict \"if-then\" rules. | Early tax software or basic chess bots. |\n| **Machine Learning (2000s)** | **Learning from Data:** Systems find patterns in historical data instead of relying on hard-coded rules. | Movie recommendations; product suggestions. |\n| **Deep Learning (2010s)** | **Neural Networks:** Huge, layered networks modeled after the brain process massive datasets. | Facial recognition; photo search. |\n| **Generative AI (2020s+)** | **Content Creation:** Models that understand context so deeply they can *generate* novel text, images, or code. | Chatbots; code assistants. |\n\n### The AI Pipeline: How It Works\n\nEvery AI system, from a simple spam filter to a self-driving car, follows a fundamental pipeline. Step through the stages below to see how an AI processes the world.\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"label\": \"1. Perceive\",\n      \"note\": \"AI receives raw input — text, audio, images, sensor data, or structured numbers. Examples: a typed sentence, a photo, a microphone recording, tabular data.\",\n      \"icon\": \"sensors\"\n    },\n    {\n      \"label\": \"2. Reason\",\n      \"note\": \"The AI processes the input using its paradigm — symbolic rules, learned patterns, or deep neural networks — to build an internal representation and infer meaning. Examples: parsing grammar, detecting edges, encoding vectors, attention over tokens.\",\n      \"icon\": \"psychology\"\n    },\n    {\n      \"label\": \"3. Act\",\n      \"note\": \"The AI produces an output based on its reasoning and confidence — a classification, a generated response, a decision, or a control signal. Examples: spam/not-spam label, a generated sentence, a steering angle, a recommendation.\",\n      \"icon\": \"rule\"\n    }\n  ]\n}\n```\n\n### Key Concepts: AI Levels\n\n| Level | Definition | Reality Today |\n|---|---|---|\n| **Narrow (Weak) AI** | Systems that excel at ONE specific, well-defined task. | Everything in production today: search ranking, speech recognition, photo tags, chatbots for narrow domains. |\n| **General (Strong) AI** | Systems that match human intelligence across ANY cognitive task. | Still research territory — no deployed system achieves this. |\n| **Superintelligence** | Systems that exceed the best human performance on every task. | Hypothetical; the subject of safety research. |\n\n| Term | Meaning | Example |\n|---|---|---|\n| **Agent** | Anything that perceives its environment and acts on it. | A self-driving car's software, a chat assistant. |\n| **Model** | A mathematical function (learned or hand-built) mapping input → output. | A spam classifier, a translation system. |\n| **Training** | The phase where the model learns patterns from data. | Showing a model millions of labeled emails. |\n| **Inference** | The phase where the trained model runs on new, unseen input. | Classifying an incoming email at runtime. |\n| **LLM (Large Language Model)** | A deep network trained on massive text to predict/manipulate language. | A chat assistant generating a response. |\n\n### Deep Learning, ML & AI: Where They Sit\n\n| Layer | Definition | What It Gives You |\n|---|---|---|\n| **AI** | The umbrella discipline: any machine that mimics intelligence. | The \"what\" — the overall goal. |\n| **Machine Learning (ML)** | A branch of AI that learns patterns from data instead of rules. | The \"how\" — a data-driven approach. |\n| **Deep Learning (DL)** | An ML subfield using multi-layer neural networks. | The \"engine\" — the technique behind modern breakthroughs. |\n\n> **The Golden Rule of AI:** \n> AI is not a single product or a magic brain. It is an overarching *discipline*. When modern tech companies talk about \"AI\" today, they are almost exclusively referring to **Machine Learning** or **Deep Learning** techniques operating beneath that umbrella.\n\n---\nNext → Let's explore how **Machine Learning** enables computers to learn from data without being explicitly programmed."},{"id":"what-is-ml","title":"What is Machine Learning (ML)?","icon":"psychology","order":2,"file":"okf/llms/what-are-llms/what-is-ml.md","body":"---\ntype: Section\ntitle: What is Machine Learning (ML)?\ndescription: AI, Machine Learning & Deep Learning - What is Machine Learning (ML)?\ntags: [what-is-ml,what-are-llms,llms]\ntimestamp: 2026-09-05T22:10:00.000Z\nsection: what-is-ml\nguide: what-are-llms\nphase: llms\nicon: psychology\norder: 2\n---\n\n# What is Machine Learning (ML)?\n\n**Icon:** psychology\n\n**Machine Learning (ML)** is a specific subfield of AI where computers learn to perform tasks without being explicitly programmed with rules. Instead of hard-coding logic (e.g., \"if an email contains the word 'lottery', mark as spam\"), you feed the computer massive amounts of **data**, and it learns to recognize patterns on its own.\n\nThink of how you teach a child to identify a cat. You don't hand them a list of rules: \"It must have pointy ears, whiskers, and a tail.\" Instead, you show them hundreds of pictures of cats and say, \"This is a cat.\" Eventually, their brain wires itself to recognize cats. ML works exactly the same way using mathematics.\n\n### The Big Four: ML Paradigms\n\nDepending on the data available, ML systems are trained in different ways. Almost every major AI product utilizes one of these paradigms:\n\n| Paradigm | How it Learns | Real-World Use Case |\n|---|---|---|\n| **Supervised** | You provide labeled examples (Input → Known Answer). The model learns the mapping. | **Email spam filter:** Trained on millions of emails manually labeled as \"Spam\" or \"Not Spam\". |\n| **Unsupervised** | You provide raw, unlabeled data. The model discovers hidden structure or groupings. | **User segmentation:** Grouping users with similar habits into hidden clusters. |\n| **Semi-supervised** | You provide a tiny bit of labeled data, and a massive amount of unlabeled data. | **Photo auto-tagging:** You label two photos of \"Mom,\" and it finds her in 5,000 unlabeled photos. |\n| **Reinforcement** | The model learns by trial and error in an environment, getting \"rewards\" or \"penalties.\" | **Game-playing agents:** An AI playing millions of games against itself to master chess or Go. |\n\n### Supervised vs. Unsupervised Learning\n\nHover over the diagrams to see the difference between feeding a model labeled answers versus letting it find structure on its own.\n\n<div class=\"my-8 w-full flex justify-center gap-8 flex-wrap\">\n  <!-- Supervised -->\n  <svg viewBox=\"0 0 300 200\" class=\"w-full max-w-xs drop-shadow-md\" xmlns=\"http://www.w3.org/2000/svg\">\n    <style>\n      .hover-scale { transition: transform 0.3s ease; transform-origin: center; cursor: pointer; }\n      .hover-scale:hover { transform: scale(1.03); }\n      .text-sm-bold { font-family: inherit; font-weight: 600; font-size: 14px; text-anchor: middle; fill: currentColor; }\n    </style>\n    <g class=\"hover-scale\">\n      <rect x=\"0\" y=\"0\" width=\"300\" height=\"200\" rx=\"16\" class=\"fill-slate-50 dark:fill-slate-800/80 stroke-slate-200 dark:stroke-slate-700\" stroke-width=\"2\" />\n      <text x=\"150\" y=\"30\" class=\"text-sm-bold theme-text\">Supervised Learning</text>\n      <!-- Labeled Data -->\n      <circle cx=\"80\" cy=\"80\" r=\"15\" class=\"fill-blue-500\" />\n      <text x=\"80\" y=\"85\" class=\"text-[10px] fill-white text-center\" text-anchor=\"middle\">Cat</text>\n      \n      <circle cx=\"80\" cy=\"130\" r=\"15\" class=\"fill-red-500\" />\n      <text x=\"80\" y=\"135\" class=\"text-[10px] fill-white text-center\" text-anchor=\"middle\">Dog</text>\n\n      <!-- Arrow -->\n      <path d=\"M 115 105 L 175 105\" stroke=\"currentColor\" stroke-width=\"2\" marker-end=\"url(#arrow)\" class=\"theme-text-muted\" />\n\n      <!-- Model Prediction -->\n      <rect x=\"190\" y=\"70\" width=\"70\" height=\"70\" rx=\"8\" class=\"fill-brand-50 dark:fill-brand-900/30 stroke-brand-500\" stroke-width=\"2\" />\n      <text x=\"225\" y=\"110\" class=\"text-[12px] font-semibold text-brand-600 dark:text-brand-400\" text-anchor=\"middle\">Maps</text>\n      <text x=\"150\" y=\"175\" class=\"text-[11px] theme-text-muted\" text-anchor=\"middle\">Learns: Data → Label</text>\n    </g>\n  </svg>\n\n  <!-- Unsupervised -->\n  <svg viewBox=\"0 0 300 200\" class=\"w-full max-w-xs drop-shadow-md\" xmlns=\"http://www.w3.org/2000/svg\">\n    <g class=\"hover-scale\">\n      <rect x=\"0\" y=\"0\" width=\"300\" height=\"200\" rx=\"16\" class=\"fill-slate-50 dark:fill-slate-800/80 stroke-slate-200 dark:stroke-slate-700\" stroke-width=\"2\" />\n      <text x=\"150\" y=\"30\" class=\"text-sm-bold theme-text\">Unsupervised Learning</text>\n      \n      <!-- Unlabeled Data scattered -->\n      <circle cx=\"60\" cy=\"70\" r=\"10\" class=\"fill-slate-400\" />\n      <circle cx=\"80\" cy=\"60\" r=\"10\" class=\"fill-slate-400\" />\n      <circle cx=\"50\" cy=\"120\" r=\"10\" class=\"fill-slate-400\" />\n      <circle cx=\"70\" cy=\"140\" r=\"10\" class=\"fill-slate-400\" />\n      \n      <!-- Arrow -->\n      <path d=\"M 115 105 L 175 105\" stroke=\"currentColor\" stroke-width=\"2\" class=\"theme-text-muted\" />\n\n      <!-- Model Clusters -->\n      <rect x=\"190\" y=\"70\" width=\"70\" height=\"70\" rx=\"8\" class=\"fill-brand-50 dark:fill-brand-900/30 stroke-brand-500\" stroke-width=\"2\" />\n      <circle cx=\"210\" cy=\"90\" r=\"8\" class=\"fill-blue-500\" />\n      <circle cx=\"225\" cy=\"85\" r=\"8\" class=\"fill-blue-500\" />\n      <circle cx=\"215\" cy=\"125\" r=\"8\" class=\"fill-red-500\" />\n      <circle cx=\"230\" cy=\"115\" r=\"8\" class=\"fill-red-500\" />\n      \n      <text x=\"150\" y=\"175\" class=\"text-[11px] theme-text-muted\" text-anchor=\"middle\">Learns: Hidden Structures</text>\n    </g>\n  </svg>\n</div>\n\n> **Garbage In, Garbage Out:** An ML model is completely dependent on its training data. If you train a resume-screening AI exclusively on resumes from men, the model will learn to downrank women. The model isn't \"smart\" enough to know right from wrong—it only memorizes the patterns you give it.\n\n### Key Concepts: Why Models Fail — Bias & Variance\n\nEvery model's error decomposes into two forces:\n\n| Error Source | Definition | Symptom | In Plain Words |\n|---|---|---|---|\n| **Bias** | Error from wrong assumptions — the model is too simple to capture the pattern. | **Underfitting:** model underperforms on BOTH training and new data. | \"The model can't fit what's there.\" |\n| **Variance** | Error from sensitivity to training data — the model memorizes noise. | **Overfitting:** model is great on training data, poor on new data. | \"The model fits too well — it memorized, not learned.\" |\n\n| | Underfitting (High Bias) | Overfitting (High Variance) |\n|---|---|---|\n| **Training error** | High | Very low (near zero) |\n| **Test/validation error** | High | High (much worse than training) |\n| **How to fix** | More features, more complex model (more layers/trees), train longer. | More data, regularization (L1/L2, dropout), simpler model, early stopping, cross-validation. |\n\n**The tradeoff:** increasing model complexity lowers bias but raises variance. The goal is the sweet spot where TOTAL error is minimized — low bias AND low variance.\n\n### Train, Validation & Test Sets\n\n| Split | Purpose | Gotcha |\n|---|---|---|\n| **Training** | Model learns parameters from this data. | The only set the model should ever \"see\" during learning. |\n| **Validation** | Tune hyperparameters & early-stop training. | Choose a split that reflects the real-world distribution. |\n| **Test (held-out)** | Final, one-time evaluation of generalization. | NEVER tune on it — once you iterate against test data, it becomes validation. |\n\n**Data leakage warning:** if any information from validation/test leaks into training (e.g., normalizing using test-set statistics, or duplicates across splits), your evaluation is fake and the model will underperform in production.\n\n### Evaluation Metrics — When Each Matters\n\n| Metric | Definition | Best When |\n|---|---|---|\n| **Accuracy** | Correct predictions / total. | Balanced classes (roughly equal examples per class). |\n| **Precision** | Of those predicted POSITIVE, how many were right? (False positives cost more.) | Spam filters, fraud alerts — you prefer fewer false alarms. |\n| **Recall** | Of all actual positives, how many did we catch? (False negatives cost more.) | Cancer screening — you cannot afford to miss a real case. |\n| **F1 Score** | Harmonic mean of precision & recall. | Imbalanced data where you need a single balance number. |\n| **ROC-AUC** | Probability that a random positive scores higher than a random negative. | Ranking tasks (search, recommendations), threshold-independent comparison. |\n\n**Imbalanced-data trap:** with 99% non-spam / 1% spam, a model that predicts \"never spam\" is 99% accurate but useless. That's exactly when precision/recall/F1 matter — accuracy lies on skewed data.\n\n### Feature Engineering — The Real Differentiator\n\n| Concept | Definition | Why It Matters |\n|---|---|---|\n| **Features** | The measurable input properties the model sees. | Garbage features → garbage model, regardless of algorithm. |\n| **Feature engineering** | Transforming raw data into useful signals by hand. | For tabular data, this usually matters MORE than the model choice. |\n| **Feature scaling** | Normalizing values (e.g., 0–1, or standard deviation). | Gradient-based models (SVM, logistic regression, neural nets) converge poorly on unscaled data. |\n| **Feature selection** | Keeping only informative features, dropping noise/duplicates. | Reduces overfitting, speeds up training, improves interpretability. |\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"label\": \"1. Collect & prepare data\",\n      \"note\": \"Gather raw examples, clean them, and split into training / validation / test sets. Examples: deduplication, handling missing values, normalizing numbers.\",\n      \"icon\": \"data_object\"\n    },\n    {\n      \"label\": \"2. Extract features\",\n      \"note\": \"Turn raw inputs into numeric features the model can operate on. Examples: pixel intensities, word counts, one-hot encoding, embeddings.\",\n      \"icon\": \"tune\"\n    },\n    {\n      \"label\": \"3. Choose a model\",\n      \"note\": \"Pick an algorithm suited to the task and data size. Examples: decision tree, logistic regression, neural network, SVM, kNN.\",\n      \"icon\": \"category\"\n    },\n    {\n      \"label\": \"4. Train\",\n      \"note\": \"Feed the model batches of data; it updates weights to reduce loss. Examples: gradient descent, backpropagation, multiple epochs, early stopping.\",\n      \"icon\": \"model_training\"\n    },\n    {\n      \"label\": \"5. Evaluate\",\n      \"note\": \"Score the trained model on held-out test data to measure generalization. Watch for overfitting: high training vs low test accuracy. Examples: accuracy, precision / recall, F1 score, ROC-AUC.\",\n      \"icon\": \"fact_check\"\n    },\n    {\n      \"label\": \"6. Deploy & infer\",\n      \"note\": \"Ship the model to production and run it on new inputs. Monitor for drift — real-world data changes over time. Examples: API endpoint, mobile app, real-time scoring, batch inference.\",\n      \"icon\": \"rocket_launch\"\n    }\n  ]\n}\n```\n\n> **The Model-Selection Rule of Thumb:** start with a simple, interpretable baseline (logistic regression or a decision tree) to establish a floor, then add complexity only if it clearly beats that floor. Simplicity wins when data is small; data volume, not algorithm exoticism, is what usually differentiates production systems.\n\n---\nNext → How ML branches into **Discriminative** vs **Generative** models."},{"id":"discriminative-vs-generative","title":"Discriminative vs Generative Models","icon":"compare","order":3,"file":"okf/llms/what-are-llms/discriminative-vs-generative.md","body":"---\ntype: Section\ntitle: Discriminative vs Generative Models\ndescription: AI, Machine Learning & Deep Learning - Discriminative vs Generative Models\ntags: [discriminative-vs-generative,what-are-llms,llms]\ntimestamp: 2026-09-05T22:15:00.000Z\nsection: discriminative-vs-generative\nguide: what-are-llms\nphase: llms\nicon: compare\norder: 3\n---\n\n# Discriminative vs Generative Models\n\n**Icon:** compare\n\nAll ML models fall into two broad families based on *what they model* about the data distribution:\n\n## Tree Data\n\n```json\n{\n  \"label\": \"Machine Learning\",\n  \"note\": \"Learn patterns from data\",\n  \"icon\": \"psychology\",\n  \"children\": [\n    {\n      \"label\": \"Discriminative Models\",\n      \"note\": \"Learn the decision boundary / P(y|x) — 'what separates classes?'\",\n      \"icon\": \"scatter_plot\",\n      \"children\": [\n        { \"label\": \"Classification\", \"note\": \"Spam vs not-spam, cat vs dog\", \"icon\": \"label\" },\n        { \"label\": \"Regression\", \"note\": \"Predict a continuous number\", \"icon\": \"trending_up\" },\n        { \"label\": \"Algorithms\", \"note\": \"Logistic regression, SVM, decision trees, CNNs (for labeling)\", \"icon\": \"schema\" }\n      ]\n    },\n    {\n      \"label\": \"Generative Models\",\n      \"note\": \"Learn the full distribution P(x) and samples — 'what does the data look like?'\",\n      \"icon\": \"auto_awesome\",\n      \"children\": [\n        { \"label\": \"Text generation\", \"note\": \"Next-token prediction in language models\", \"icon\": \"text_fields\" },\n        { \"label\": \"Image generation\", \"note\": \"Diffusion, GANs, VAEs\", \"icon\": \"image\" },\n        { \"label\": \"Audio / video\", \"note\": \"Speech synthesis, deepfakes\", \"icon\": \"graphic_eq\" }\n      ]\n    }\n  ]\n}\n```\n\n| Dimension | Discriminative | Generative |\n|---|---|---|\n| **Learns** | Boundary between classes | The whole data distribution |\n| **Maps** | Input → label (`P(y|x)`) | Input → distribution (`P(x)`) |\n| **Output** | A classification / number | New, synthetic data |\n| **Typical use** | Spam detection, fraud, ranking | Chat, art, code, speech |\n| **Data need** | Efficient with less data | Hungry for massive data |\n| **Main risk** | Can't create anything new | Hallucination, cost |\n\n**Simple intuition:** a **discriminative** model learns where to *draw the line* between cats and dogs so it can classify a new photo. A **generative** model learns what cats and dogs *generally look like* so it can draw a brand-new cat.\n\n**Golden rule:** use **discriminative** for fast, accurate decision/classification on fixed inputs; use **generative** when you need to *create* new content. LLMs (large language models) are generative models.\n\n---\nNext → Discriminative Algorithms (Logistic Regression, SVM, Decision Trees, kNN, LDA, boosted trees)."},{"id":"discriminative-algorithms","title":"Discriminative Algorithms","icon":"scatter_plot","order":4,"file":"okf/llms/what-are-llms/discriminative-algorithms.md","body":"---\ntype: Section\ntitle: Discriminative Algorithms\ndescription: AI, Machine Learning & Deep Learning - Discriminative Algorithms (Logistic Regression, SVM, Decision Trees, kNN, LDA, Boosted Trees, DNN)\ntags: [discriminative-algorithms,what-are-llms,llms]\ntimestamp: 2026-09-05T22:25:00.000Z\nsection: discriminative-algorithms\nguide: what-are-llms\nphase: llms\nicon: scatter_plot\norder: 4\n---\n\n# Discriminative Algorithms\n\n**Icon:** scatter_plot\n\nDiscriminative models learn the conditional probability **P(y|x)** — the probability of a label given an input. They directly model the **decision boundary** that separates classes. For regression, they learn **P(y|x)** directly as a function f(x).\n\n**Core idea:** instead of learning how data is generated, discriminative models learn *where to draw the line* between classes. This makes them more sample-efficient for classification tasks.\n\n## Tree Data\n\n```json\n{\n  \"label\": \"Discriminative Models (P(y|x))\",\n  \"note\": \"Learn decision boundaries directly\",\n  \"icon\": \"scatter_plot\",\n  \"children\": [\n    {\n      \"label\": \"Linear Models\",\n      \"note\": \"Linear decision boundaries\",\n      \"icon\": \"linear_scale\",\n      \"children\": [\n        { \"label\": \"Logistic Regression\", \"note\": \"Binary & multi-class classification via sigmoid\", \"icon\": \"functions\" },\n        { \"label\": \"Linear Discriminant Analysis (LDA)\", \"note\": \"Gaussian class-conditional with shared covariance\", \"icon\": \"analytics\" }\n      ]\n    },\n    {\n      \"label\": \"Margin-based\",\n      \"note\": \"Maximize margin between classes\",\n      \"icon\": \"tune\",\n      \"children\": [\n        { \"label\": \"SVM (Linear / RBF Kernel)\", \"note\": \"Max-margin classifier; kernel trick for non-linear\", \"icon\": \"graphic_eq\" }\n      ]\n    },\n    {\n      \"label\": \"Tree-based\",\n      \"note\": \"Axis-aligned splits, interpretable\",\n      \"icon\": \"account_tree\",\n      \"children\": [\n        { \"label\": \"Decision Tree\", \"note\": \"Recursive partitioning by information gain/Gini\", \"icon\": \"call_split\" },\n        { \"label\": \"Random Forest\", \"note\": \"Bagged trees + feature subsampling\", \"icon\": \"forest\" },\n        { \"label\": \"Gradient Boosted Trees (XGBoost/LightGBM/CatBoost)\", \"note\": \"Sequential boosting with gradient descent in function space\", \"icon\": \"trending_up\" }\n      ]\n    },\n    {\n      \"label\": \"Instance-based\",\n      \"note\": \"Memorize training data, local decisions\",\n      \"icon\": \"near_me\",\n      \"children\": [\n        { \"label\": \"k-Nearest Neighbors (kNN)\", \"note\": \"Majority vote of k closest neighbors\", \"icon\": \"scatter_plot\" }\n      ]\n    },\n    {\n      \"label\": \"Neural (Discriminative)\",\n      \"note\": \"Deep non-linear decision boundaries\",\n      \"icon\": \"hub\",\n      \"children\": [\n        { \"label\": \"MLP / DNN\", \"note\": \"Fully-connected layers for tabular classification\", \"icon\": \"account_tree\" },\n        { \"label\": \"CNN (for classification)\", \"note\": \"Spatial hierarchies for images\", \"icon\": \"imagesearch_roller\" },\n        { \"label\": \"Transformer (encoder)\", \"note\": \"Self-attention for sequence classification\", \"icon\": \"bolt\" }\n      ]\n    }\n  ]\n}\n```\n\n---\n\n### FAANG-Ready Cheat Cards\n\n| Algorithm | Key Idea | Probability Model | Loss | Complexity (Train / Infer) | When to Use | Interview One-Liner |\n|---|---|---|---|---|---|---|\n| **Logistic Regression** | Linear decision boundary + sigmoid | `P(y=1|x) = σ(wᵀx + b)` | Binary Cross-Entropy | O(nd) / O(d) | Baseline, interpretable, small/medium data | \"Logistic regression = linear regression + sigmoid; it's a linear classifier on P(y\\|x).\" |\n| **Linear SVM** | Max-margin hyperplane | Deterministic: `sign(wᵀx + b)` | Hinge Loss | O(n²d) to O(n³) / O(d) | High-dim sparse data (text), clear margin | \"SVM finds the widest street separating classes; support vectors define the boundary.\" |\n| **Kernel SVM (RBF)** | Non-linear via kernel trick | Implicit high-dim mapping | Hinge + Kernel | O(n²d) to O(n³) / O(n_sv) | Complex boundaries, medium data | \"Kernel SVM = linear SVM in infinite dimensions without computing them explicitly.\" |\n| **Decision Tree** | Recursive axis-aligned splits | Piecewise constant P(y|x) | Gini / Entropy | O(n d log n) / O(depth) | Interpretable, mixed features, non-linear | \"Trees partition feature space into rectangles; depth controls overfitting.\" |\n| **Random Forest** | Bagging + feature sampling | Avg of tree posteriors | — | O(B n d log n) / O(B·depth) | Robust baseline, feature importance | \"Random Forest = many decorrelated trees voted; reduces variance without increasing bias.\" |\n| **Gradient Boosted Trees** | Additive model: fit residuals | Sequential gradient descent | Any diff loss | O(B n d) / O(B·depth) | Tabular SOTA, winning competitions | \"GBDT fits trees to negative gradients of loss; shrinkage + subsampling = regularization.\" |\n| **k-Nearest Neighbors** | Local majority vote | Non-parametric: empirical P(y|x) | — | O(1) / O(n d) (or KD-tree) | Small data, low-dim, baseline | \"kNN has no training — all work at inference; curse of dimensionality kills it in high-d.\" |\n| **LDA** | Gaussian classes, shared Σ | `P(x|y) ~ N(μ_y, Σ)`, Bayes rule | MLE on Σ, μ | O(n d²) / O(d²) | Small data, Gaussian-ish classes | \"LDA = generative assumptions (Gaussian + shared cov) → linear boundary; related to logistic.\" |\n| **MLP / DNN** | Compose non-linear layers | `P(y|x) = softmax(W_L · relu(...))` | Cross-Entropy | O(n d w L) / O(d w L) | Complex patterns, large data | \"MLP = universal approximator; depth = hierarchical features; ReLU avoids vanishing gradients.\" |\n\n---\n\n**Golden rule:** Start with Logistic Regression or Random Forest as baselines. Use SVM for high-dim sparse data. Use GBDT (XGBoost/LightGBM) for tabular SOTA. Use DNNs when data is massive and features are unstructured (images, text, audio).\n\n---\nNext → Generative Algorithms (Naive Bayes, GMM, VAE, GAN, Diffusion, Autoregressive)."},{"id":"generative-algorithms","title":"Generative Algorithms","icon":"auto_awesome","order":5,"file":"okf/llms/what-are-llms/generative-algorithms.md","body":"---\ntype: Section\ntitle: Generative Algorithms\ndescription: AI, Machine Learning & Deep Learning - Generative Algorithms (Naive Bayes, GMM, HMM, VAE, GAN, Diffusion, Autoregressive)\ntags: [generative-algorithms,what-are-llms,llms]\ntimestamp: 2026-09-05T22:30:00.000Z\nsection: generative-algorithms\nguide: what-are-llms\nphase: llms\nicon: auto_awesome\norder: 5\n---\n\n# Generative Algorithms\n\n**Icon:** auto_awesome\n\nGenerative models learn the joint distribution **P(x, y)** or the marginal **P(x)**. They model *how data is generated* — enabling sampling of new, synthetic examples. For classification, they apply Bayes' rule: `P(y|x) = P(x|y)P(y) / P(x)`.\n\n**Core idea:** instead of learning a decision boundary, generative models learn *what the data looks like* so they can create new instances that look like the training data.\n\n## Tree Data\n\n```json\n{\n  \"label\": \"Generative Models (P(x) / P(x,y))\",\n  \"note\": \"Learn data distribution, then sample\",\n  \"icon\": \"auto_awesome\",\n  \"children\": [\n    {\n      \"label\": \"Explicit Density (Exact Likelihood)\",\n      \"note\": \"Tractable P(x) — exact likelihood training\",\n      \"icon\": \"functions\",\n      \"children\": [\n        { \"label\": \"Naive Bayes\", \"note\": \"P(x|y) = ∏ P(x_i|y) — feature independence\", \"icon\": \"category\" },\n        { \"label\": \"Gaussian Mixture Model (GMM)\", \"note\": \"P(x) = ∑ π_k N(x|μ_k, Σ_k) — soft clustering\", \"icon\": \"scatter_plot\" },\n        { \"label\": \"Hidden Markov Model (HMM)\", \"note\": \"Sequential P(x_1...x_T|y) with latent states\", \"icon\": \"timeline\" },\n        { \"label\": \"Linear Discriminant Analysis (LDA)\", \"note\": \"Gaussian P(x|y) with shared Σ → also generative\", \"icon\": \"analytics\" }\n      ]\n    },\n    {\n      \"label\": \"Implicit Density (Sampling, No Exact P(x))\",\n      \"note\": \"Learn to sample without computing P(x) explicitly\",\n      \"icon\": \"hub\",\n      \"children\": [\n        { \"label\": \"Variational Autoencoder (VAE)\", \"note\": \"ELBO maximization; latent z ~ N; reparameterization trick\", \"icon\": \"account_tree\" },\n        { \"label\": \"Generative Adversarial Network (GAN)\", \"note\": \"Min-max game: Generator vs Discriminator\", \"icon\": \"sports_esports\" },\n        { \"label\": \"Diffusion / Score-based\", \"note\": \"Denoise score matching; reverse SDE/ODE\", \"icon\": \"auto_fix_high\" }\n      ]\n    },\n    {\n      \"label\": \"Autoregressive (Factorized Likelihood)\",\n      \"note\": \"P(x) = ∏ P(x_t | x_<t) — sequential generation\",\n      \"icon\": \"bolt\",\n      \"children\": [\n        { \"label\": \"PixelCNN / PixelRNN\", \"note\": \"Autoregressive over pixels; masked convolutions\", \"icon\": \"imagesearch_roller\" },\n        { \"label\": \"WaveNet\", \"note\": \"Autoregressive audio; dilated convolutions\", \"icon\": \"graphic_eq\" },\n        { \"label\": \"GPT / Transformer Decoder\", \"note\": \"Next-token prediction; causal self-attention\", \"icon\": \"text_fields\" }\n      ]\n    }\n  ]\n}\n```\n\n---\n\n### FAANG-Ready Cheat Cards\n\n| Algorithm | Key Idea | Probability Model | Training Objective | Sample Quality | When to Use | Interview One-Liner |\n|---|---|---|---|---|---|---|\n| **Naive Bayes** | Feature independence `P(x|y)=∏P(x_i|y)` | `P(y|x) ∝ P(y)∏P(x_i|y)` | MLE on P(x_i|y), P(y) | Low (strong assumptions) | Text classification, spam, tiny data | \"Naive Bayes = Bayes rule + conditional independence; fast, works surprisingly well for text.\" |\n| **Gaussian Mixture Model (GMM)** | Mixture of K Gaussians | `P(x) = ∑ π_k N(x\\|μ_k, Σ_k)` | EM algorithm (E-step: γ, M-step: μ,Σ,π) | Medium (ellipsoidal clusters) | Soft clustering, density estimation | \"GMM = soft k-means; EM alternates responsibility assignment and parameter updates.\" |\n| **Hidden Markov Model (HMM)** | Latent discrete states, Markov transitions | `P(x_1:T) = ∑ P(z_1:T)∏P(x_t\\|z_t)P(z_t\\|z_{t-1})` | Baum-Welch (EM) / Viterbi (decode) | Sequence modeling | Speech, POS tagging, bioinformatics | \"HMM = Markov chain on steroids — latent states generate observations.\" |\n| **LDA (Generative)** | Gaussian per class, shared Σ | `P(x\\|y) ~ N(μ_y, Σ)`, `P(y)=π_y` | MLE on Σ, μ_y, π_y | Linear boundaries | When Gaussian assumption holds, small data | \"LDA's generative view explains why it gives linear boundaries — equal covariance assumption.\" |\n| **VAE** | Latent z ~ N; maximize ELBO | `P(x) = ∫ P(x\\|z)P(z)dz ≈ ELBO` | ELBO = E[log P(x\\|z)] - KL[q(z\\|x)\\|P(z)] | Blurry (pixel-wise MSE) | Anomaly detection, representation learning | \"VAE = encoder q(z\\|x) + decoder P(x\\|z); reparameterization trick enables backprop through sampling.\" |\n| **GAN** | Min-max: G fools D, D catches G | Implicit P_G(x) = G(z), z~P(z) | min_G max_D V(D,G) = E[log D(x)] + E[log(1-D(G(z)))] | Sharp (adversarial) | Image synthesis, style transfer | \"GAN = counterfeiter (G) vs detective (D); Nash equilibrium = P_G = P_data.\" |\n| **Diffusion** | Learn reverse denoising process | `P(x_0) = ∫ P(x_T) ∏ P(x_{t-1}\\|x_t) dx_T` | Score matching: ∥ε_θ(x_t,t) - ε∥² | SOTA (images, audio) | Image gen (Stable Diffusion), audio | \"Diffusion = destroy with noise, learn to reverse; score function guides denoising.\" |\n| **Autoregressive (GPT)** | Next-token P(x_t\\|x_<t) | `P(x) = ∏ P(x_t\\|x_<t)` | Cross-entropy on next token | Coherent, long-range | LLMs (GPT, LLaMA), code gen | \"GPT = Transformer decoder + causal mask + massive scale; in-context learning emerges.\" |\n\n---\n\n**Golden rule:** Naive Bayes/GMM for classical baselines & interpretability. VAE for latent representations & anomaly detection. GAN for sharp images (but training is finicky). Diffusion for SOTA generation quality. Autoregressive Transformers (GPT) for sequential data — text, code, audio — where coherence over long range matters.\n\n---\nNext → Deep Learning & its Algorithms (MLP, CNN, RNN/LSTM, Transformer, Backprop, Optimizers)."},{"id":"what-is-deep-learning","title":"What is Deep Learning (DL)?","icon":"hub","order":6,"file":"okf/llms/what-are-llms/what-is-deep-learning.md","body":"---\ntype: Section\ntitle: What is Deep Learning (DL)?\ndescription: AI, Machine Learning & Deep Learning - What is Deep Learning (DL)?\ntags: [what-is-deep-learning,what-are-llms,llms]\ntimestamp: 2026-09-05T22:20:00.000Z\nsection: what-is-deep-learning\nguide: what-are-llms\nphase: llms\nicon: hub\norder: 6\n---\n\n# What is Deep Learning (DL)?\n\n**Icon:** hub\n\n**Deep Learning (DL)** is a highly advanced subfield of Machine Learning powered by **Artificial Neural Networks**. The word \"deep\" simply refers to the number of layers in the network—the more layers, the deeper the network, and the more complex patterns it can understand.\n\nThink of Deep Learning like a massive, highly-organized factory assembly line. If you want to recognize a face:\n1. The **first layer** of workers only looks for basic edges and lines.\n2. The **second layer** combines those edges to find shapes like circles (eyes) or triangles (noses).\n3. The **third layer** combines the shapes to identify the entire face.\n\nInstead of engineers manually telling the computer what features to look for (like in traditional ML), a deep network **learns its own features** automatically from raw data.\n\n### Why Deep Learning Runs the World\n\nDeep Learning is the engine behind almost every modern AI breakthrough because it scales incredibly well with data and compute power.\n\n| Architecture | What it Excels At | Real-World Use Case |\n|---|---|---|\n| **CNN (Convolutional)** | Computer Vision (Images/Video) | **Facial recognition:** Recognizing spatial patterns in a 3D scan of a face. |\n| **RNN (Recurrent)** | Sequential Data (Time-series/Speech) | **Speech processing:** Transcribing audio over time. |\n| **Transformer** | Context & Language at massive scale | **Text generation:** Generating human-like text by attending to massive amounts of context simultaneously. |\n\n### Key Concepts: Why Deep Learning, Why Now\n\nDeep Learning's dominance rests on a three-legged stool:\n\n| Leg | What Changed | Result |\n|---|---|---|\n| **Data** | The internet generated massive labeled/unlabeled datasets. | Networks finally had enough examples to learn rich patterns. |\n| **Compute** | GPUs/Tensor Cores made matrix math thousands of times faster. | Training deep networks became practical instead of decades-long. |\n| **Algorithms** | Better activations (ReLU), initialization, optimization (Adam), normalization (BatchNorm). | Training deep networks stopped failing from vanishing gradients. |\n\nTake away any leg and modern deep learning collapses — which is why \"deep learning\" research didn't go mainstream until ~2012.\n\n### Deep Learning vs Classical Machine Learning\n\n| Dimension | Classical ML (trees, SVM, logistic regression) | Deep Learning |\n|---|---|---|\n| **Feature handling** | Requires hand-engineered features. | Learns features automatically from raw data. |\n| **Best data type** | Tabular / structured data (spreadsheets, SQL). | Unstructured data (images, text, audio). |\n| **Data needed** | Works with thousands of rows. | Often needs millions of examples. |\n| **Compute** | Trains on a laptop CPU. | Needs GPUs / TPUs for real models. |\n| **Interpretability** | High (trees, linear models). | Low (\"black box\") — needs tools like attention maps / LIME / SHAP. |\n| **When to use** | Small/medium tabular data, constrained compute, need to explain predictions. | Huge unstructured datasets, SOTA accuracy is the goal, GPU available. |\n\n**Rule of thumb:** for a spreadsheet, start classical (gradient-boosted trees usually win); for images/audio/text, deep learning is the default.\n\n### The Neural Network: Layers of Abstraction\n\nHover over the neural network below. Watch how raw data (input) gets passed through hidden layers (the \"factory workers\"), where each layer extracts deeper meaning before making a final prediction.\n\n<div class=\"my-8 w-full flex justify-center\">\n  <svg viewBox=\"0 0 600 250\" class=\"w-full max-w-2xl drop-shadow-md\" xmlns=\"http://www.w3.org/2000/svg\">\n    <style>\n      .node { fill: currentColor; transition: transform 0.2s ease, fill 0.2s ease; cursor: pointer; }\n      .node:hover { transform: scale(1.3); }\n      .edge { stroke: currentColor; opacity: 0.3; }\n      @keyframes pulse { 0% { opacity: 0.3; } 50% { opacity: 0.8; } 100% { opacity: 0.3; } }\n      .active-edge { animation: pulse 1.5s infinite; stroke: #3b82f6; opacity: 0.8; stroke-width: 2; }\n      .label-text { font-family: inherit; font-size: 12px; font-weight: 600; text-anchor: middle; fill: currentColor; }\n    </style>\n\n    <!-- Connections Input -> Hidden 1 -->\n    <line x1=\"100\" y1=\"75\" x2=\"250\" y2=\"50\" class=\"edge theme-text-muted active-edge\" />\n    <line x1=\"100\" y1=\"125\" x2=\"250\" y2=\"100\" class=\"edge theme-text-muted active-edge\" style=\"animation-delay: 0.2s;\" />\n    <line x1=\"100\" y1=\"175\" x2=\"250\" y2=\"150\" class=\"edge theme-text-muted active-edge\" style=\"animation-delay: 0.4s;\" />\n    \n    <!-- Connections Hidden 1 -> Hidden 2 -->\n    <line x1=\"250\" y1=\"50\" x2=\"400\" y2=\"75\" class=\"edge theme-text-muted\" />\n    <line x1=\"250\" y1=\"100\" x2=\"400\" y2=\"125\" class=\"edge theme-text-muted\" />\n    <line x1=\"250\" y1=\"150\" x2=\"400\" y2=\"175\" class=\"edge theme-text-muted\" />\n    <line x1=\"250\" y1=\"200\" x2=\"400\" y2=\"125\" class=\"edge theme-text-muted\" />\n    \n    <!-- Connections Hidden 2 -> Output -->\n    <line x1=\"400\" y1=\"75\" x2=\"550\" y2=\"125\" class=\"edge theme-text-muted active-edge\" style=\"animation-delay: 0.6s;\" />\n    <line x1=\"400\" y1=\"125\" x2=\"550\" y2=\"125\" class=\"edge theme-text-muted active-edge\" style=\"animation-delay: 0.8s;\" />\n    <line x1=\"400\" y1=\"175\" x2=\"550\" y2=\"125\" class=\"edge theme-text-muted active-edge\" style=\"animation-delay: 1.0s;\" />\n\n    <!-- Input Layer -->\n    <circle cx=\"100\" cy=\"75\" r=\"15\" class=\"node theme-bg-subtle text-slate-400\" />\n    <circle cx=\"100\" cy=\"125\" r=\"15\" class=\"node theme-bg-subtle text-slate-400\" />\n    <circle cx=\"100\" cy=\"175\" r=\"15\" class=\"node theme-bg-subtle text-slate-400\" />\n    <text x=\"100\" y=\"230\" class=\"label-text theme-text\">Input Layer</text>\n    <text x=\"100\" y=\"245\" class=\"label-text theme-text-muted text-[9px] font-normal\">Raw Pixels / Words</text>\n\n    <!-- Hidden Layer 1 -->\n    <circle cx=\"250\" cy=\"50\" r=\"15\" class=\"node text-brand-400\" />\n    <circle cx=\"250\" cy=\"100\" r=\"15\" class=\"node text-brand-400\" />\n    <circle cx=\"250\" cy=\"150\" r=\"15\" class=\"node text-brand-400\" />\n    <circle cx=\"250\" cy=\"200\" r=\"15\" class=\"node text-brand-400\" />\n    <text x=\"250\" y=\"230\" class=\"label-text text-brand-500\">Hidden Layer 1</text>\n    <text x=\"250\" y=\"245\" class=\"label-text theme-text-muted text-[9px] font-normal\">Finds Edges & Lines</text>\n\n    <!-- Hidden Layer 2 -->\n    <circle cx=\"400\" cy=\"75\" r=\"15\" class=\"node text-brand-600\" />\n    <circle cx=\"400\" cy=\"125\" r=\"15\" class=\"node text-brand-600\" />\n    <circle cx=\"400\" cy=\"175\" r=\"15\" class=\"node text-brand-600\" />\n    <text x=\"400\" y=\"230\" class=\"label-text text-brand-600\">Hidden Layer 2</text>\n    <text x=\"400\" y=\"245\" class=\"label-text theme-text-muted text-[9px] font-normal\">Finds Shapes</text>\n\n    <!-- Output Layer -->\n    <circle cx=\"550\" cy=\"125\" r=\"15\" class=\"node theme-bg text-emerald-500 border border-emerald-500\" />\n    <text x=\"550\" y=\"230\" class=\"label-text theme-text\">Output Layer</text>\n    <text x=\"550\" y=\"245\" class=\"label-text theme-text-muted text-[9px] font-normal\">Final Prediction</text>\n  </svg>\n</div>\n\n### Activation Functions\n\nAn activation function decides how much a neuron \"fires\" by introducing non-linearity — without it, stacking layers would collapse into a single linear operation.\n\n| Activation | Formula | Strengths | Failure Mode |\n|---|---|---|---|\n| **Sigmoid** | σ(x) = 1 / (1 + e⁻ˣ) | Squashes to (0,1) — good for probabilities. | Vanishing gradient at extremes; not zero-centered. |\n| **Tanh** | tanh(x) | Zero-centered, outputs (-1,1). | Still vanishes at extremes. |\n| **ReLU** | max(0, x) | Cheap, non-saturating — default for hidden layers. | **Dead ReLU:** neurons can permanently output 0. Fix: Leaky ReLU. |\n| **GELU** | x · Φ(x) | Smooth ReLU variant; used in modern transformers. | Slightly more compute than ReLU. |\n| **Softmax** | e^zᵢ / Σe^zⱼ | Output layer for multi-class → probabilities summing to 1. | Not a hidden-layer choice; only for the final layer. |\n\n**Mental model for backpropagation:** the network computes predictions with a **forward pass**, then compares to the true answer to compute loss; **backpropagation** walks the error backward through the layers using the **chain rule**, computing how much each weight contributed, and updates each weight by that amount scaled by the **learning rate**. Repeat over batches until loss stops improving.\n\n> **The Hardware Catch:** Deep learning's superpower is that its accuracy keeps improving as you throw more data at it. However, doing this requires millions of complex mathematical operations, which is why Deep Learning is heavily dependent on extremely powerful hardware like GPUs (Graphics Processing Units).\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"label\": \"1. Input layer\",\n      \"note\": \"Raw data enters as numeric vectors — pixels, words, audio samples. Examples: image pixels, embedded tokens, sensor values.\",\n      \"icon\": \"keyboard_input\"\n    },\n    {\n      \"label\": \"2. Hidden layers\",\n      \"note\": \"Each layer transforms the data, learning progressively abstract features. Examples: edges → shapes → objects, phrases → meaning, frequencies → phonemes.\",\n      \"icon\": \"layers\"\n    },\n    {\n      \"label\": \"3. Activation\",\n      \"note\": \"A non-linear function decides how much a neuron 'fires', adding expressive power. Examples: ReLU, Sigmoid, Tanh, GELU, Swish.\",\n      \"icon\": \"bolt\"\n    },\n    {\n      \"label\": \"4. Output layer\",\n      \"note\": \"Produces the final prediction — a class, a number, or a distribution over next tokens. Examples: softmax probabilities, regression value, logits.\",\n      \"icon\": \"output\"\n    },\n    {\n      \"label\": \"5. Loss computation\",\n      \"note\": \"Compare the prediction to the true answer and compute an error score. Examples: cross-entropy, mean squared error, hinge loss.\",\n      \"icon\": \"error\"\n    },\n    {\n      \"label\": \"6. Backpropagation\",\n      \"note\": \"Propagate the error backwards and update every weight to lower the loss. Examples: gradient descent, Adam optimizer, learning rate scheduling.\",\n      \"icon\": \"settings_backup_restore\"\n    }\n  ]\n}\n```\n\n> **Detecting overfitting during training:** after every epoch, compare training loss to validation loss. Training keeps dropping while validation plateaus or rises → the network is memorizing, not learning. That's when you add regularization (dropout, weight decay), gather more data, or shrink the model.\n\n---\n**You've now completed the foundation:** AI (the field) → ML (learn from data) → Discriminative vs Generative (what a model learns) → Discriminative Algorithms → Generative Algorithms → Deep Learning (the neural-network engine). Modern LLMs are built on top of this — decoder transformers that are deep, generative models."}]},{"id":"inference","title":"Inference","description":"What inference is, how LLM inference works (prefill vs decode), and an interactive architecture pipeline diagram.","icon":"bolt","order":2,"sections":[{"id":"what-is-inference","title":"What is an Inference?","icon":"bolt","order":1,"file":"okf/llms/inference/what-is-inference.md","body":"---\ntype: Section\ntitle: What is an Inference?\ndescription: Inference - What is an Inference?\ntags: [what-is-inference,inference,llms]\ntimestamp: 2026-08-26T23:11:16.578Z\nsection: what-is-inference\nguide: inference\nphase: llms\nicon: bolt\norder: 1\n---\n\n# What is an Inference?\n\n**Icon:** bolt\n\n**Inference** is the phase where a trained model actually *produces output* from new input — as opposed to **training**, where the model *learns* weights. For an LLM, inference means: take a prompt, run it through the Transformer, and generate tokens.\n\n**Training vs inference:**\n\n| | Training | Inference |\n|---|---|---|\n| Goal | Learn weights from data | Produce output from input |\n| Runs | Once, offline, expensive | Repeatedly, per request |\n| Needs | Labels, gradients, backward pass | Only forward pass |\n| Memory | Gradients + optimizer state | Activations + KV cache |\n| Output | Updated model | Generated tokens |\n\n**Why it matters:** inference is where the model meets users. Its cost, latency, and throughput — not training — determine the real-world bill and experience. Everything from batching to quantization exists to make inference cheaper and faster.\n\n**Golden rule:** training builds the model; inference *runs* it. Optimizing inference (not training) is what makes an LLM usable in production.\n"},{"id":"how-it-works","title":"How it works?","icon":"sync","order":2,"file":"okf/llms/inference/how-it-works.md","body":"---\ntype: Section\ntitle: How it works?\ndescription: Inference - How it works?\ntags: [how-it-works,inference,llms]\ntimestamp: 2026-08-26T23:11:16.578Z\nsection: how-it-works\nguide: inference\nphase: llms\nicon: sync\norder: 2\n---\n\n# How it works?\n\n**Icon:** sync\n\nLLM inference is a two-phase loop driven by the **Transformer forward pass**:\n\n**1. Prefill (prompt processing):** the whole input prompt is fed through the model in parallel (one batched forward pass). The model computes the Key/Value vectors for every prompt token and stores them in the **KV cache**. Output: the logits for the first generated token.\n\n**2. Decode (token generation):** the model generates **one token at a time**. Each step:\n- reads the *last* token + the cached K/V of all previous tokens,\n- computes new K/V (cached),\n- produces logits → samples the next token,\n- appends it and repeats until an end-of-sequence token or max length.\n\n**Prefill vs decode at a glance:**\n\n| | Prefill | Decode |\n|---|---|---|\n| Compute | Parallel over all prompt tokens | Sequential, one token per step |\n| KV cache | Filled here | Read + extended here |\n| Bottleneck | Compute-bound (big matmuls) | Memory-bandwidth-bound (reads weights + cache) |\n\n**Batching:** in production, many requests are batched together (continuous / batch scheduling) so the GPU stays busy during both phases. This is the single biggest throughput lever.\n\n**Golden rule:** prefill is compute-bound, decode is memory-bound. Good serving stacks optimize each phase separately and batch across requests.\n"},{"id":"inference-architecture","title":"Architecture (Interactive Pipeline)","icon":"account_tree","order":3,"file":"okf/llms/inference/inference-architecture.md","body":"---\ntype: Section\ntitle: Architecture (Interactive Pipeline)\ndescription: Inference - Architecture (Interactive Pipeline)\ntags: [inference-architecture,inference,llms]\ntimestamp: 2026-08-26T23:11:16.578Z\nsection: inference-architecture\nguide: inference\nphase: llms\nicon: account_tree\norder: 3\n---\n\n# Architecture (Interactive Pipeline)\n\n**Icon:** account_tree\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"icon\": \"text_fields\",\n      \"label\": \"Input Prompt\",\n      \"note\": \"Raw, untokenized user text enters the system. For example, a prompt like Translate to French: Hello is still just a string of characters here, with no tokenization yet.\"\n    },\n    {\n      \"icon\": \"token\",\n      \"label\": \"Tokenizer\",\n      \"note\": \"Splits text into subword tokens (e.g. BPE). Each token maps to an integer ID. Roughly 1 token ≈ 4 characters of English.\"\n    },\n    {\n      \"icon\": \"grid_on\",\n      \"label\": \"Embedding\",\n      \"note\": \"Token IDs are mapped to dense vectors that capture meaning. Positional encodings are added so token order is preserved.\"\n    },\n    {\n      \"icon\": \"account_tree\",\n      \"label\": \"Transformer (Prefill + Decode)\",\n      \"note\": \"Prefill: process the whole prompt in parallel and fill the KV cache. Decode: generate one token at a time, reusing the cache (see the KV Cache guide).\"\n    },\n    {\n      \"icon\": \"functions\",\n      \"label\": \"LM Head / Logits\",\n      \"note\": \"The final layer outputs a logit vector over the vocabulary — a raw score for every possible next token.\"\n    },\n    {\n      \"icon\": \"tune\",\n      \"label\": \"Sampling\",\n      \"note\": \"Logits become probabilities via softmax, then a token is chosen: greedy (argmax), or temperature / top-p / top-k for controlled diversity.\"\n    },\n    {\n      \"icon\": \"output\",\n      \"label\": \"Output Token → Loop\",\n      \"note\": \"The chosen token is emitted, appended to the sequence, and fed back in for the next decode step until an end token or max length.\"\n    }\n  ]\n}\n```\n"}]},{"id":"prompt","title":"Prompt","description":"What a prompt is, the main types of prompts, the architecture of a well-structured prompt, and concrete examples.","icon":"edit_note","order":3,"sections":[{"id":"what-is-prompt","title":"What is a Prompt?","icon":"chat","order":1,"file":"okf/llms/prompt/what-is-prompt.md","body":"---\ntype: Section\ntitle: What is a Prompt?\ndescription: Prompt - What is a Prompt?\ntags: [what-is-prompt,prompt,llms]\ntimestamp: 2026-08-26T23:11:16.579Z\nsection: what-is-prompt\nguide: prompt\nphase: llms\nicon: chat\norder: 1\n---\n\n# What is a Prompt?\n\n**Icon:** chat\n\nA **prompt** is the input you give an LLM to elicit a desired response — natural-language (and sometimes structured) instructions, context, and examples. Prompting is how you *steer* a model without retraining: the model's behavior is largely determined by what you put in the prompt.\n\n**Why prompting matters:** an LLM is a frozen, next-token predictor. The prompt is the only live lever you have at inference time to control task, format, tone, and correctness. Small prompt changes can mean the difference between a useless answer and a great one.\n\n**Key concepts:**\n- **Prompt ≠ fine-tuning:** prompting changes input, not weights.\n- **Context window:** everything you put in consumes tokens (and memory via the KV cache).\n- **Determinism vs sampling:** same prompt + low temperature → stable output; high temperature → varied.\n\n**Golden rule:** a prompt is a contract with the model. Be explicit about role, task, format, and constraints, and the model will meet you halfway.\n"},{"id":"types-of-prompts","title":"Types of Prompts","icon":"category","order":2,"file":"okf/llms/prompt/types-of-prompts.md","body":"---\ntype: Section\ntitle: Types of Prompts\ndescription: Prompt - Types of Prompts\ntags: [types-of-prompts,prompt,llms]\ntimestamp: 2026-08-26T23:11:16.579Z\nsection: types-of-prompts\nguide: prompt\nphase: llms\nicon: category\norder: 2\n---\n\n# Types of Prompts\n\n**Icon:** category\n\nPrompts come in families. Know the main ones:\n\n| Type | What it is | When to use |\n|---|---|---|\n| **Zero-shot** | Ask directly, no examples | Simple, well-defined tasks |\n| **Few-shot** | Give 1+ input→output examples | When format/style must be demonstrated |\n| **System / Role** | Set persona + rules up front | Define behavior across a session |\n| **Instruction** | Step-by-step commands | Multi-step or precise tasks |\n| **Chain-of-Thought** | \"Think step by step\" | Reasoning, math, logic |\n| **Contextual / RAG** | Inject retrieved docs | Grounding on private/live data |\n| **Negative** | State what NOT to do | Avoid known failure modes |\n\n**Zero-shot vs few-shot:** zero-shot relies on the model's priors; few-shot *shows* the pattern, which dramatically improves consistency on structured or unusual tasks.\n\n**Golden rule:** start zero-shot, add few-shot examples when the format is fragile, and use Chain-of-Thought only when reasoning is the bottleneck.\n"},{"id":"system-vs-user-prompt","title":"System vs User Prompt","icon":"swap_horiz","order":3,"file":"okf/llms/prompt/system-vs-user-prompt.md","body":"---\ntype: Section\ntitle: System vs User Prompt\ndescription: Prompt - System vs User Prompt\ntags: [system-vs-user-prompt,prompt,llms]\ntimestamp: 2026-08-26T23:11:16.579Z\nsection: system-vs-user-prompt\nguide: prompt\nphase: llms\nicon: swap_horiz\norder: 3\n---\n\n# System vs User Prompt\n\n**Icon:** swap_horiz\n\nIn chat-style LLM APIs, a prompt is usually split into **message roles**. The two you interact with most are the **system prompt** and the **user prompt** (alongside the assistant's own replies).\n\n| | System prompt | User prompt |\n|---|---|---|\n| Set by | The developer / app | The end user |\n| When | Once, at the start of a session | Every turn |\n| Purpose | Persona, rules, global behavior | The actual question or task |\n| Persistence | Stays for the whole conversation | Varies per message |\n| Example | \"You are a concise SQL expert.\" | \"List the top 5 customers by revenue.\" |\n\n**System prompt:** developer-supplied instructions that shape *how* the model behaves for the entire session — tone, persona, guardrails, output format. It is typically not shown to the end user.\n\n**User prompt:** the input the person types each turn — the question, command, or data the model should act on.\n\n**Why the split matters:** keeping behavior (system) separate from content (user) makes the same assistant reusable across users and tasks, and lets you update guardrails without rewriting every request.\n\n**Golden rule:** put stable behavior and constraints in the system prompt; put the task and data in the user prompt.\n"},{"id":"prompt-architecture","title":"Prompt Architecture (Interactive)","icon":"account_tree","order":4,"file":"okf/llms/prompt/prompt-architecture.md","body":"---\ntype: Section\ntitle: Prompt Architecture (Interactive)\ndescription: Prompt - Prompt Architecture (Interactive)\ntags: [prompt-architecture,prompt,llms]\ntimestamp: 2026-08-26T23:11:16.579Z\nsection: prompt-architecture\nguide: prompt\nphase: llms\nicon: account_tree\norder: 4\n---\n\n# Prompt Architecture (Interactive)\n\n**Icon:** account_tree\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"icon\": \"shield_person\",\n      \"label\": \"System Message\",\n      \"note\": \"Sets the model role/persona and global rules (tone, safety, constraints). Persists across the conversation.\"\n    },\n    {\n      \"icon\": \"history\",\n      \"label\": \"Context\",\n      \"note\": \"Retrieved documents or prior conversation turns (RAG). Grounds the answer in real, current data.\"\n    },\n    {\n      \"icon\": \"list_alt\",\n      \"label\": \"Instruction\",\n      \"note\": \"The actual task: what to do, step by step. The clearest instruction wins even on a weak model.\"\n    },\n    {\n      \"icon\": \"format_quote\",\n      \"label\": \"Few-shot Examples\",\n      \"note\": \"Demonstration input→output pairs that show the desired pattern and format. Powerful for structured tasks.\"\n    },\n    {\n      \"icon\": \"chat\",\n      \"label\": \"User Input\",\n      \"note\": \"The live query. Combined with everything above, this is what the model actually responds to.\"\n    },\n    {\n      \"icon\": \"code\",\n      \"label\": \"Output Format\",\n      \"note\": \"Constraints on the response: JSON schema, length, style, or citation rules. Forces machine-readable, parseable output.\"\n    }\n  ]\n}\n```\n"},{"id":"examples-of-prompts","title":"Examples of Prompts","icon":"apps","order":5,"file":"okf/llms/prompt/examples-of-prompts.md","body":"---\ntype: Section\ntitle: Examples of Prompts\ndescription: Prompt - Examples of Prompts\ntags: [examples-of-prompts,prompt,llms]\ntimestamp: 2026-08-26T23:11:16.580Z\nsection: examples-of-prompts\nguide: prompt\nphase: llms\nicon: apps\norder: 5\n---\n\n# Examples of Prompts\n\n**Icon:** apps\n\n**1. Zero-shot classification**\n> Classify the sentiment of this review as Positive, Negative, or Neutral: \"The battery died after two weeks.\" → Answer:\n\n**2. Few-shot extraction (JSON)**\n> Extract entities as JSON.\n> Example: \"Apple bought a startup in Seattle.\" → {\"org\":\"Apple\",\"city\":\"Seattle\"}\n> Text: \"OpenAI hired a researcher from London.\" →\n\n**3. Chain-of-Thought math**\n> A train travels 60 km in 45 minutes. What is its speed in km/h? Think step by step.\n\n**4. RAG grounded answer**\n> Using the provided policy document, answer: what is the refund window? Only use the document.\n\n**5. System + negative constraint**\n> You are a senior Rust reviewer. Explain the bug. Do NOT rewrite the whole file; only show the minimal fix.\n\n**Golden rule:** show, don't just tell — few-shot examples and explicit output formats beat long paragraphs of instructions.\n"},{"id":"prompt-compression-optimization","title":"Compression & Optimization","icon":"compress","order":6,"file":"okf/llms/prompt/prompt-compression-optimization.md","body":"---\ntype: Section\ntitle: Compression & Optimization\ndescription: Prompt - Compression & Optimization\ntags: [prompt-compression-optimization,prompt,llms]\ntimestamp: 2026-08-26T23:11:16.580Z\nsection: prompt-compression-optimization\nguide: prompt\nphase: llms\nicon: compress\norder: 6\n---\n\n# Compression & Optimization\n\n**Icon:** compress\n\nAs prompts grow, they cost tokens, latency, and context-window space — and the KV cache (see the KV Cache guide) grows with every token. **Compression** shrinks the prompt; **optimization** makes it cheaper and more reliable.\n\n**Compression techniques:**\n- **Truncation / windowing:** keep only the most recent N turns or characters.\n- **Summarization:** compress old context into a rolling summary instead of raw history.\n- **Selective context:** retrieve only the documents or facts actually needed (RAG), not the whole corpus.\n- **Semantic compression:** replace verbose text with dense embeddings or learned summaries the model can expand.\n- **Prompt caching:** mark the stable prefix (system prompt, few-shot examples) as cacheable so repeated calls reuse the KV cache instead of recomputing it.\n\n**Optimization techniques:**\n- **Few-shot pruning:** keep only the examples that actually move the output; drop the rest.\n- **Instruction tightening:** shorter, explicit instructions beat long paragraphs.\n- **Template & variable reuse:** fixed templates plus minimal per-request variables reduce tokens and drift.\n- **Deferred / lazy context:** load heavy context only when the task needs it.\n- **Batching:** group independent prompts to amortize model overhead.\n\n**Others to know:**\n- **Evaluation & A/B testing:** measure task success before and after changes — optimize what you can measure.\n- **Versioning:** treat prompts like code; track changes and roll back.\n- **Guardrails & validation:** validate the output (schema, filters) rather than stuffing more instructions into the prompt.\n\n**Golden rule:** shrink what is stable (and cache it), retrieve only what is needed (do not dump everything), and measure — most \"better prompting\" is really better compression and better evaluation.\n"}]},{"id":"forward-propagation","title":"Forward Propagation","description":"What forward propagation is, how it works in neural networks and transformers, an interactive architecture diagram, and the difference from backward propagation.","icon":"forward","order":4,"sections":[{"id":"what-is-forward-propagation","title":"What is Forward Propagation?","icon":"forward","order":1,"file":"okf/llms/forward-propagation/what-is-forward-propagation.md","body":"---\ntype: Section\ntitle: What is Forward Propagation?\ndescription: Forward Propagation - What is Forward Propagation?\ntags: [what-is-forward-propagation,forward-propagation,llms]\ntimestamp: 2026-08-26T23:11:16.580Z\nsection: what-is-forward-propagation\nguide: forward-propagation\nphase: llms\nicon: forward\norder: 1\n---\n\n# What is Forward Propagation?\n\n**Icon:** forward\n\nA **forward pass** (or forward propagation) is the computation of output data through a neural network when given an input. It proceeds layer by layer — from input to output — applying each layer's weights and activation functions, but **never updating any parameter**: it only computes predictions.\n\n**Why it matters:**\n- **Inference is 100% forward.** Every time you query an LLM, it runs a forward pass over the prompt (plus cached KV states for prior tokens).\n- **Training needs two phases:** first forward (to get the loss) then backward (to compute gradients). Forward alone does nothing to learn.\n\n> In a transformer:  \n> Input → Embedding → Attention → FFN → LayerNorm → Output\n\n**Golden rule:** forward propagation is a *deterministic pipeline*. Same input + same weights = same output. No learning happens during the forward pass itself.\n"},{"id":"how-forward-works","title":"How Forward Propagation Works","icon":"sync","order":2,"file":"okf/llms/forward-propagation/how-forward-works.md","body":"---\ntype: Section\ntitle: How Forward Propagation Works\ndescription: Forward Propagation - How Forward Propagation Works\ntags: [how-forward-works,forward-propagation,llms]\ntimestamp: 2026-08-26T23:11:16.580Z\nsection: how-forward-works\nguide: forward-propagation\nphase: llms\nicon: sync\norder: 2\n---\n\n# How Forward Propagation Works\n\n**Icon:** sync\n\nA single layer transforms its input as:\n\n> z = W × x + b      (linear: weight matrix × input + bias)  \n> a = activation(z)  (non-linear: ReLU, GELU, softmax)\n\nAfter the first layer's activation becomes the input to the second layer, and so on until the final output.\n\n**In a Transformer block** (the architecture used by every modern LLM), one forward step does:\n\n1. **Multi-Head Attention:** compute queries (Q), keys (K), values (V) for all tokens, apply scaled dot-product attention, then concatenate heads. Output = attention weights × V.\n2. **Add & Norm:** add the residual (original input) to the attention output, then apply layer normalization.\n3. **Feed-Forward Network (FFN):** apply a small MLP — up-project (W₁), non-linearity (GELU), down-project (W₂). Position-wise: each token transformed independently.\n4. **Add & Norm:** another residual connection + normalization.\n\n**Stacking layers:** N identical blocks are stacked. The output of block L feeds block L+1. This depth is what gives transformers their representational power.\n\n**Cost:** each layer performs O(seq_len² × d) for attention (with seq_len = token count, d = hidden size) and O(seq_len × d × 4) for the feed-forward. Why compute-bound prefill vs memory-bound decode.\n"},{"id":"forward-architecture-diagram","title":"Forward Architecture (Interactive)","icon":"account_tree","order":3,"file":"okf/llms/forward-propagation/forward-architecture-diagram.md","body":"---\ntype: Section\ntitle: Forward Architecture (Interactive)\ndescription: Forward Propagation - Forward Architecture (Interactive)\ntags: [forward-architecture-diagram,forward-propagation,llms]\ntimestamp: 2026-08-26T23:11:16.581Z\nsection: forward-architecture-diagram\nguide: forward-propagation\nphase: llms\nicon: account_tree\norder: 3\n---\n\n# Forward Architecture (Interactive)\n\n**Icon:** account_tree\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"icon\": \"text_fields\",\n      \"label\": \"Input Tokens\",\n      \"note\": \"Token IDs and positions enter. For inference: prompt tokens plus any cached KV history.\"\n    },\n    {\n      \"icon\": \"grid_on\",\n      \"label\": \"Embedding + Position\",\n      \"note\": \"Token IDs become dense vectors. Positional encodings inject order so the model knows sequence.\"\n    },\n    {\n      \"icon\": \"mode_comment\",\n      \"label\": \"Attention Layer\",\n      \"note\": \"Q, K, V projections → scaled dot-product → weighted values. Output combines information from other tokens.\"\n    },\n    {\n      \"icon\": \"add\",\n      \"label\": \"Add & Norm\",\n      \"note\": \"Residual connection adds input back (gradient flow). LayerNorm stabilizes deep stacking.\"\n    },\n    {\n      \"icon\": \"layers\",\n      \"label\": \"Feed-Forward (FFN)\",\n      \"note\": \"Up-project (d → 4d), GELU non-linearity, down-project (4d → d). Position-wise MLP per token.\"\n    },\n    {\n      \"icon\": \"add\",\n      \"label\": \"Add & Norm\",\n      \"note\": \"Second residual + norm after FFN. Completes one Transformer block. Repeat N times.\"\n    },\n    {\n      \"icon\": \"functions\",\n      \"label\": \"Output Logits\",\n      \"note\": \"Final linear projection to vocabulary-sized logits — the raw scores for the next token.\"\n    }\n  ]\n}\n```\n"},{"id":"forward-vs-backpropagation","title":"Forward vs Backprop","icon":"swap_vert","order":4,"file":"okf/llms/forward-propagation/forward-vs-backpropagation.md","body":"---\ntype: Section\ntitle: Forward vs Backprop\ndescription: Forward Propagation - Forward vs Backprop\ntags: [forward-vs-backpropagation,forward-propagation,llms]\ntimestamp: 2026-08-26T23:11:16.581Z\nsection: forward-vs-backpropagation\nguide: forward-propagation\nphase: llms\nicon: swap_vert\norder: 4\n---\n\n# Forward vs Backprop\n\n**Icon:** swap_vert\n\nA simple comparison of the two core phases in training:\n\n| | Forward Propagation | Backward Propagation |\n|---|---|---|\n| Direction | Input → output | Output → input |\n| Computes | Predictions, loss | Gradients of loss w.r.t. weights |\n| Updates weights? | No | Yes (via optimizer) |\n| Runs during | Inference AND training | Training only |\n| Relative cost | ~1× | ~2–3× (forward + backward) |\n\n**Why you mostly care about forward:** in production you only run forward passes. Optimizing forward latency, memory, and KV cache reuse (see the Inference guide) is what makes LLMs fast and cheap.\n\n**Key insight:** the weights learned during training define the *forward function* that gives the model its abilities. Inference just applies that function.\n\n**Golden rule:** forward = predict, backward = learn. Ship fast forward passes; training happens offline.\n"}]},{"id":"backward-propagation","title":"Backward Propagation","description":"What backward propagation is, how it works through the chain rule, and its role in training neural networks.","icon":"arrow_back","order":5,"sections":[{"id":"what-is-backward-propagation","title":"What is Backward Propagation?","icon":"arrow_back","order":1,"file":"okf/llms/backward-propagation/what-is-backward-propagation.md","body":"---\ntype: Section\ntitle: What is Backward Propagation?\ndescription: Backward Propagation - What is Backward Propagation?\ntags: [what-is-backward-propagation,backward-propagation,llms]\ntimestamp: 2026-08-26T23:11:16.581Z\nsection: what-is-backward-propagation\nguide: backward-propagation\nphase: llms\nicon: arrow_back\norder: 1\n---\n\n# What is Backward Propagation?\n\n**Icon:** arrow_back\n\nBackward propagation (backprop) is the algorithm that computes gradients of the loss function with respect to every weight in the neural network. It runs *after* the forward pass and enables the model to learn by updating its weights via gradient descent.\n\n**Why it exists:** an LLM (or any neural net) has billions of parameters. To make those parameters useful, they must be updated based on how wrong the forward pass was. Backprop tells us precisely how much to change each weight.\n\n**Key facts:**\n- **Efficient via dynamic programming:** chains gradients backward through the computational graph, reusing intermediate results (no recomputation).\n- **Chain rule everywhere:** the gradient of a composite function is the product of gradients at each step. This is why we store activations during the forward pass.\n- **Works for any differentiable function:** ReLU, GELU, softmax, attention — all have known gradients.\n\n**Golden rule:** you can only backpropagate through operations whose gradients are defined (no non-differentiable jumps like if-statements on values).\n\n**Why the term \"backprop\":** gradients flow backward through the network (from loss to input layer), but the weight updates improve the forward pass for future data.\n"},{"id":"how-backward-works","title":"How Backward Propagation Works","icon":"sync","order":2,"file":"okf/llms/backward-propagation/how-backward-works.md","body":"---\ntype: Section\ntitle: How Backward Propagation Works\ndescription: Backward Propagation - How Backward Propagation Works\ntags: [how-backward-works,backward-propagation,llms]\ntimestamp: 2026-08-26T23:11:16.581Z\nsection: how-backward-works\nguide: backward-propagation\nphase: llms\nicon: sync\norder: 2\n---\n\n# How Backward Propagation Works\n\n**Icon:** sync\n\nGiven a loss L (e.g., cross-entropy), backprop computes ∂L/∂θ for every weight θ in the network.\n\n**The chain rule in action:**\nFor a multi-layer network, denote Layer ℓ's output as aℓ, weight as Wℓ, activation as σ (e.g., ReLU):\n\n> zℓ = Wℓ · aℓ₋₁ + bℓ                (linear forward)\n> aℓ = σ(zℓ)                           (activation)\n\nThe gradient w.r.t. aℓ is:\n> δℓ = ∂L/∂aℓ = δℓ₊₁  Wℓ₊₁ᵀ  σ'(zℓ)    (propagate from next layer)\n\nThen the gradients for the weights:\n> ∂L/∂Wℓ = δℓ  aℓ₋₁ᵀ                   (outer product)\n\n**Backprop algorithm (high level):**\n1. **Forward pass:** compute all activations a₀, a₁, ..., aₙ; compute loss L.\n2. **Backward pass:** compute δₙ = ∂L/∂aₙ, then δℓ = δℓ₊₁ Wℓ₊₁ᵀ σ'(zℓ) for ℓ = n-1 down to 1.\n3. **Gradient accumulation:** ∂L/∂Wℓ = δℓ aℓ₋₁ᵀ.\n4. **Update:** Wℓ ← Wℓ - η ∂L/∂Wℓ (gradient descent step, η = learning rate).\n\n**In a Transformer block:**\n- Attention: gradients flow through softmax and dot-product; must remember QKᵀ scale.\n- KV cache: during backprop, gradients accumulate for all positions (no caching benefits).\n- The backward pass is why Transformers are memory-heavy: we need to store all activations for the gradient computation.\n"},{"id":"backward-architecture-diagram","title":"Backward Architecture (Interactive)","icon":"account_tree","order":3,"file":"okf/llms/backward-propagation/backward-architecture-diagram.md","body":"---\ntype: Section\ntitle: Backward Architecture (Interactive)\ndescription: Backward Propagation - Backward Architecture (Interactive)\ntags: [backward-architecture-diagram,backward-propagation,llms]\ntimestamp: 2026-08-26T23:11:16.581Z\nsection: backward-architecture-diagram\nguide: backward-propagation\nphase: llms\nicon: account_tree\norder: 3\n---\n\n# Backward Architecture (Interactive)\n\n**Icon:** account_tree\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"icon\": \"functions\",\n      \"label\": \"Loss L\",\n      \"note\": \"The objective function (e.g., cross-entropy). Higher loss = more wrong predictions.\"\n    },\n    {\n      \"icon\": \"sync\",\n      \"label\": \"Upstream Gradient\",\n      \"note\": \"∂L/∂logits flows backward from the loss into the network.\"\n    },\n    {\n      \"icon\": \"layers\",\n      \"label\": \"Layer N (FFN)\",\n      \"note\": \"Apply chain rule: d = d_past * W^T * ReLU(z). Compute gradients for weights and bias.\"\n    },\n    {\n      \"icon\": \"add\",\n      \"label\": \"Add & Norm\",\n      \"note\": \"Gradients split: one for residual path, one for LayerNorm. Sum and normalize.\"\n    },\n    {\n      \"icon\": \"mode_comment\",\n      \"label\": \"Attention Block\",\n      \"note\": \"Backprop through softmax (softmax × (input - sumsoftmax)). Compute QK/V gradients.\"\n    },\n    {\n      \"icon\": \"grid_on\",\n      \"label\": \"Embedding\",\n      \"note\": \"Embedding gradients sum over all token positions they appear in. Rare tokens get bigger updates.\"\n    },\n    {\n      \"icon\": \"text_fields\",\n      \"label\": \"Input Tokens\",\n      \"note\": \"Final gradient w.r.t. input tokens. Used in adversarial training, gradient-based attacks, or input embedding analysis.\"\n    }\n  ]\n}\n```\n"}]},{"id":"transformers","title":"Transformers","description":"What the Transformer architecture is, how attention works, and why it dominates modern LLMs.","icon":"account_tree","order":6,"sections":[{"id":"what-is-transformer","title":"What is a Transformer?","icon":"account_tree","order":1,"file":"okf/llms/transformers/what-is-transformer.md","body":"---\ntype: Section\ntitle: What is a Transformer?\ndescription: Transformers - What is a Transformer?\ntags: [what-is-transformer,transformers,llms]\ntimestamp: 2026-08-26T23:11:16.582Z\nsection: what-is-transformer\nguide: transformers\nphase: llms\nicon: account_tree\norder: 1\n---\n\n# What is a Transformer?\n\n**Icon:** account_tree\n\nA **Transformer** is a neural network architecture introduced in 2017 (\"Attention Is All You Need\") that relies entirely on **self-attention** mechanisms to model relationships between tokens in a sequence. Unlike recurrent networks (RNNs, LSTMs), Transformers process all tokens in parallel, making them dramatically faster and more scalable.\n\n**Why it matters:** virtually every modern LLM — GPT, Claude, Llama, Gemini, Mistral — is built on the Transformer architecture. Understanding Transformers means understanding how today's AI actually works.\n\n**Key characteristics:**\n- **Parallel processing:** all tokens processed simultaneously during training (not sequentially like RNNs).\n- **Self-attention:** each token attends to every other token, capturing context regardless of distance.\n- **Multi-head attention:** multiple attention heads run in parallel, each learning different relationships (syntax, semantics, coreference).\n- **Positional encoding:** since there's no recurrence, position information is injected via sinusoidal or learned encodings.\n- **Scalable:** the architecture scales well with more data, more parameters, and longer sequences.\n\n**Golden rule:** Transformers are not \"intelligent\" — they are sophisticated pattern matchers that use attention to weigh the importance of every token in context. Their power comes from scale, not from understanding.\n"},{"id":"how-attention-works","title":"How Self-Attention Works","icon":"sync","order":2,"file":"okf/llms/transformers/how-attention-works.md","body":"---\ntype: Section\ntitle: How Self-Attention Works\ndescription: Transformers - How Self-Attention Works\ntags: [how-attention-works,transformers,llms]\ntimestamp: 2026-08-26T23:11:16.582Z\nsection: how-attention-works\nguide: transformers\nphase: llms\nicon: sync\norder: 2\n---\n\n# How Self-Attention Works\n\n**Icon:** sync\n\nSelf-attention is the core mechanism of the Transformer. It allows each token to dynamically focus on relevant tokens in the sequence.\n\n**The three vectors:**\nFor each token, the model computes three vectors:\n- **Query (Q):** \"What am I looking for?\"\n- **Key (K):** \"What do I contain?\"\n- **Value (V):** \"What information do I offer?\"\n\n**Attention formula:**\n> Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V\n\nWhere:\n- **QKᵀ:** measures similarity between every query and every key\n- **√dₖ:** scaling factor (prevents dot products from growing too large)\n- **softmax:** converts similarities to probabilities (weights)\n- **V:** weighted sum of values (the actual output)\n\n**Step by step:**\n1. Compute Q, K, V for every token via learned linear projections.\n2. Compute attention scores: Q × Kᵀ for every pair of tokens.\n3. Scale scores by 1/√dₖ.\n4. Apply softmax to get attention weights (probabilities).\n5. Multiply weights by V and sum — each token's output is a weighted mix of all tokens' values.\n\n**Multi-head attention:** instead of one attention function, run h parallel \"heads\" with different learned projections, then concatenate and project. Each head can learn different relationships — one might track subject-verb agreement, another might resolve pronouns.\n\n**Golden rule:** attention is a *weighted average*. Every token's output is a blend of information from the entire sequence, weighted by relevance. The model learns what to attend to.\n"},{"id":"transformer-architecture-diagram","title":"Transformer Architecture (Interactive)","icon":"account_tree","order":3,"file":"okf/llms/transformers/transformer-architecture-diagram.md","body":"---\ntype: Section\ntitle: Transformer Architecture (Interactive)\ndescription: Transformers - Transformer Architecture (Interactive)\ntags: [transformer-architecture-diagram,transformers,llms]\ntimestamp: 2026-08-26T23:11:16.583Z\nsection: transformer-architecture-diagram\nguide: transformers\nphase: llms\nicon: account_tree\norder: 3\n---\n\n# Transformer Architecture (Interactive)\n\n**Icon:** account_tree\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"icon\": \"text_fields\",\n      \"label\": \"Input Tokens\",\n      \"note\": \"Raw text is tokenized into integer IDs. Each token becomes a dense vector via embedding.\"\n    },\n    {\n      \"icon\": \"grid_on\",\n      \"label\": \"Token + Position Embedding\",\n      \"note\": \"Token embeddings capture meaning; positional encodings inject order information. Since Transformers have no recurrence, position is essential.\"\n    },\n    {\n      \"icon\": \"mode_comment\",\n      \"label\": \"Multi-Head Self-Attention\",\n      \"note\": \"Each token computes Q, K, V and attends to every other token. Multiple heads run in parallel, each learning different relationships (syntax, semantics, coreference).\"\n    },\n    {\n      \"icon\": \"add\",\n      \"label\": \"Add & Norm (Residual)\",\n      \"note\": \"Attention output is added to the original input (residual connection), then normalized. This stabilizes training and allows gradients to flow through deep stacks.\"\n    },\n    {\n      \"icon\": \"layers\",\n      \"label\": \"Feed-Forward Network (FFN)\",\n      \"note\": \"Position-wise MLP: up-project to 4× hidden size, apply GELU non-linearity, down-project back. Each token transformed independently — this is where \\\"knowledge\\\" is stored.\"\n    },\n    {\n      \"icon\": \"add\",\n      \"label\": \"Add & Norm (Residual)\",\n      \"note\": \"Second residual + LayerNorm after FFN. One complete Transformer block.\"\n    },\n    {\n      \"icon\": \"repeat\",\n      \"label\": \"Repeat N times\",\n      \"note\": \"Typical LLMs stack 12–96+ identical blocks. Deeper stacking = more capacity. The output of block L feeds block L+1.\"\n    },\n    {\n      \"icon\": \"functions\",\n      \"label\": \"Final LayerNorm + LM Head\",\n      \"note\": \"Final normalization, then linear projection to vocabulary size. Output: logits — raw scores for every possible next token.\"\n    },\n    {\n      \"icon\": \"tune\",\n      \"label\": \"Sampling → Output\",\n      \"note\": \"Convert logits to probabilities via softmax, then sample (greedy, temperature, top-p, top-k). Emit token, append to sequence, repeat for autoregressive generation.\"\n    }\n  ]\n}\n```\n"},{"id":"encoder-vs-decoder","title":"Encoder vs Decoder Architectures","icon":"compare","order":4,"file":"okf/llms/transformers/encoder-vs-decoder.md","body":"---\ntype: Section\ntitle: Encoder vs Decoder Architectures\ndescription: Transformers - Encoder vs Decoder Architectures\ntags: [encoder-vs-decoder,transformers,llms]\ntimestamp: 2026-08-26T23:11:16.583Z\nsection: encoder-vs-decoder\nguide: transformers\nphase: llms\nicon: compare\norder: 4\n---\n\n# Encoder vs Decoder Architectures\n\n**Icon:** compare\n\nTransformers come in three architectural flavors, each suited to different tasks:\n\n**1. Encoder-only (e.g., BERT, RoBERTa)**\n- Processes input bidirectionally — each token attends to ALL other tokens (both left and right).\n- No generation capability (no autoregressive decoding).\n- **Best for:** classification, sentiment analysis, named entity recognition, search ranking.\n- **How it works:** input → Transformer encoder stack → [CLS] token or pooled output → task head.\n\n**2. Decoder-only (e.g., GPT, Llama, Claude)**\n- Processes tokens left-to-right (causal attention). Each token can only attend to previous tokens.\n- Autoregressive: generates one token at a time, feeding output back as input.\n- **Best for:** text generation, dialogue, code completion, open-ended tasks.\n- **How it works:** input → Transformer decoder stack → logits → sample next token → append → repeat.\n\n**3. Encoder-Decoder (e.g., T5, BART)**\n- Full encoder (bidirectional) processes input, then full decoder (causal) generates output.\n- Cross-attention connects encoder output to decoder layers.\n- **Best for:** translation, summarization, structured input→output tasks.\n- **How it works:** input → encoder → context vectors → decoder (with cross-attention) → output tokens.\n\n**Golden rule:** decoder-only for generation, encoder-only for understanding, encoder-decoder when you need both. Most modern LLMs use decoder-only because generation is the dominant use case.\n"},{"id":"why-transformers-win","title":"Why Transformers Dominate","icon":"trending_up","order":5,"file":"okf/llms/transformers/why-transformers-win.md","body":"---\ntype: Section\ntitle: Why Transformers Dominate\ndescription: Transformers - Why Transformers Dominate\ntags: [why-transformers-win,transformers,llms]\ntimestamp: 2026-08-26T23:11:16.583Z\nsection: why-transformers-win\nguide: transformers\nphase: llms\nicon: trending_up\norder: 5\n---\n\n# Why Transformers Dominate\n\n**Icon:** trending_up\n\nBefore Transformers (pre-2017), the dominant architectures were RNNs, LSTMs, and GRUs. Transformers displaced them for several reasons:\n\n**1. Parallelization**\n- RNNs process tokens sequentially — token t must finish before token t+1 begins.\n- Transformers process all tokens simultaneously. Training is dramatically faster on GPUs/TPUs.\n\n**2. Long-range dependencies**\n- RNNs struggle to connect distant tokens (vanishing gradients).\n- Every token attends directly to every other token in O(1) path length.\n\n**3. Scalability**\n- The architecture is simple and uniform — just stacked attention + FFN blocks.\n- Scales well with more data, more parameters, and longer sequences (with optimizations).\n\n**4. Flexibility**\n- Pre-train once on massive text, then fine-tune or prompt for any task.\n- The same architecture handles classification, generation, translation, summarization, and more.\n\n**The trade-offs:**\n- **Quadratic complexity:** self-attention is O(n²) in sequence length. Long sequences are expensive (mitigated by KV cache, FlashAttention, sliding window, etc.).\n- **Memory:** storing all activations during training is memory-heavy (see Backward Propagation guide).\n- **Data hungry:** requires massive datasets to reach peak performance.\n\n**Golden rule:** Transformers win because they are parallel, scalable, and flexible. Their quadratic attention cost is the main limitation — which is why research focuses on efficient attention, KV caching, and alternative architectures.\n"}]},{"id":"caching","title":"Caching","description":"What caching is, why it matters, and the types of caching in LLMs (KV cache, prompt cache, model cache).","icon":"memory","order":7,"sections":[{"id":"what-is-caching","title":"What is Caching?","icon":"memory","order":1,"file":"okf/llms/caching/what-is-caching.md","body":"---\ntype: Section\ntitle: What is Caching?\ndescription: Caching - What is Caching?\ntags: [what-is-caching,caching,llms]\ntimestamp: 2026-08-26T23:11:16.583Z\nsection: what-is-caching\nguide: caching\nphase: llms\nicon: memory\norder: 1\n---\n\n# What is Caching?\n\n**Icon:** memory\n\nCaching is the storage of intermediate computations or results to avoid redundant work. In LLMs, caching is crucial for performance and cost optimization.\n\n**Why caching matters:**\n- **Speed:** skip repeated computation, get instant results on cache hits.\n- **Cost:** fewer GPU cycles = lower electricity and compute dollars.\n- **Scalability:** enables serving many concurrent requests efficiently.\n\n**Golden rule:** cache what is expensive to compute and stable across requests. Never cache what changes frequently.\n\n**Types of caching in LLMs:**\n- **KV cache:** keys/values from attention layers during generation (the most impactful).\n- **Prompt cache:** embeddings of static prompt components.\n- **Model cache:** intermediate layers or weights for fast inference.\n- **Tokenizer cache:** tokenization results for repeated text segments.\n"},{"id":"kv-cache-overview","title":"KV Cache Overview","icon":"layers","order":2,"file":"okf/llms/caching/kv-cache-overview.md","body":"---\ntype: Section\ntitle: KV Cache Overview\ndescription: Caching - KV Cache Overview\ntags: [kv-cache-overview,caching,llms]\ntimestamp: 2026-08-26T23:11:16.583Z\nsection: kv-cache-overview\nguide: caching\nphase: llms\nicon: layers\norder: 2\n---\n\n# KV Cache Overview\n\n**Icon:** layers\n\nThe **KV cache** (Key-Value cache) is the most critical caching mechanism in Transformers for autoregressive generation. It stores the Q, K, and V projections computed during the attention operation.\n\n**How it works:**\n1. During prefill (prompt processing), compute K and V for each prompt token.\n2. During decode (generation), compute K and V for each new token.\n3. All K and V pairs are stored in the cache for future steps.\n\n**Attention step with cache:**\n> Output = softmax(Q · Kᵀ / √dₖ) · V\n> Where K and V come from the cache, not recomputed.\n\n**KV cache size:**\n> cache_size = 2 (K + V) × sequence_length × num_layers × num_heads × head_dim\n\n**What gets cached:**\n- Keys and values are the *heavy* parts to compute (O(n²) attention complexity).\n- They are *stable* across generation steps (once computed, they never change).\n- Embedding lookups and FFN computations are cheaper (O(n) per token).\n\n**Golden rule:** the KV cache trades *memory* for *compute*. It's why decoding a 1000-token response uses far less compute than recomputing everything from scratch each step.\n"},{"id":"caching-architecture","title":"Caching Architecture (Interactive)","icon":"account_tree","order":3,"file":"okf/llms/caching/caching-architecture.md","body":"---\ntype: Section\ntitle: Caching Architecture (Interactive)\ndescription: Caching - Caching Architecture (Interactive)\ntags: [caching-architecture,caching,llms]\ntimestamp: 2026-08-26T23:11:16.583Z\nsection: caching-architecture\nguide: caching\nphase: llms\nicon: account_tree\norder: 3\n---\n\n# Caching Architecture (Interactive)\n\n**Icon:** account_tree\n\n## Pipeline Diagram\n\n```json\n{\n  \"stages\": [\n    {\n      \"icon\": \"text_fields\",\n      \"label\": \"Input Tokens\",\n      \"note\": \"Raw text is tokenized. For each token: compute embedding + position encoding.\"\n    },\n    {\n      \"icon\": \"mode_comment\",\n      \"label\": \"Compute Q, K, V\",\n      \"note\": \"Projection matrices W_Q, W_K, W_V applied to embeddings. Expensive matrix multiplies (O(n²)).\"\n    },\n    {\n      \"icon\": \"memory\",\n      \"label\": \"KV Cache Storage\",\n      \"note\": \"Store computed K and V pairs per layer. Grows linearly with sequence length, quadratically with model size.\"\n    },\n    {\n      \"icon\": \"repeat\",\n      \"label\": \"Prefill Phase\",\n      \"note\": \"All prompt tokens processed in one batched forward pass. Fill entire KV cache up front. Heavy compute, one-time cost.\"\n    },\n    {\n      \"icon\": \"repeat\",\n      \"label\": \"Decode Phase\",\n      \"note\": \"Each generation step: compute fresh K,V for NEW token → append to cache. Subsequent steps read from cache only.\"\n    },\n    {\n      \"icon\": \"bolt\",\n      \"label\": \"Attention with Cache\",\n      \"note\": \"QKᵀ uses cached keys; softmax weighted sum uses cached values. No recomputation of prompt token K/V pairs.\"\n    },\n    {\n      \"icon\": \"functions\",\n      \"label\": \"Output + Cache Growth\",\n      \"note\": \"Generate next token → append its K,V to cache. Cache size increases by one slot per generated token.\"\n    }\n  ]\n}\n```\n"},{"id":"cache-strategies","title":"Caching Strategies","icon":"layers_clear","order":4,"file":"okf/llms/caching/cache-strategies.md","body":"---\ntype: Section\ntitle: Caching Strategies\ndescription: Caching - Caching Strategies\ntags: [cache-strategies,caching,llms]\ntimestamp: 2026-08-26T23:11:16.583Z\nsection: cache-strategies\nguide: caching\nphase: llms\nicon: layers_clear\norder: 4\n---\n\n# Caching Strategies\n\n**Icon:** layers_clear\n\nDifferent caching approaches optimize for speed, memory, or cost:\n\n**1. Full KV Cache**\n- Store ALL keys and values for the entire sequence.\n- **Pros:** fastest generation, simplest implementation.\n- **Cons:** memory-heavy; O(n) memory growth per generated token.\n\n**2. Sliding Window Cache**\n- Keep only the most recent N tokens in cache.\n- **Pros:** constant memory regardless of sequence length.\n- **Cons:** cannot attend to tokens older than window (loss of long context).\n\n**3. Block-wise / Paged Cache**\n- Partition cache into blocks; evict least-recently-used.\n- **Pros:** fine-grained control over memory budget.\n- **Cons:** more complex eviction logic, higher overhead.\n\n**4. Grouped-Query Attention (GQA)**\n- Share keys/values across multiple attention heads.\n- **Pros:** reduces KV cache memory by factor of heads.\n- **Cons:** trades off expressivity; heads compete for same information.\n\n**5. Multi-Head Latent Attention (MLA)**\n- Compress K and V representations; store compressed versions.\n- **Pros:** drastically smaller cache footprint.\n- **Cons:** extra compression/decompression cost; approximation error.\n\n**Golden rule:** choose cache strategy based on your use case:\n- **Real-time chat:** full KV cache (speed is king).\n- **Long-context summarization:** sliding window + attention mechanisms.\n- **Cost-sensitive deployment:** GQA/MLA with careful tuning.\n- **Research/innovation:** experiment with new cache eviction policies.\n"},{"id":"caching-optimization","title":"Caching Optimization","icon":"tune","order":5,"file":"okf/llms/caching/caching-optimization.md","body":"---\ntype: Section\ntitle: Caching Optimization\ndescription: Caching - Caching Optimization\ntags: [caching-optimization,caching,llms]\ntimestamp: 2026-08-26T23:11:16.584Z\nsection: caching-optimization\nguide: caching\nphase: llms\nicon: tune\norder: 5\n---\n\n# Caching Optimization\n\n**Icon:** tune\n\nCaching works, but optimizations can make it dramatically more efficient:\n\n**1. Prompt Caching**\n- Cache embeddings of static prompt components (system prompt, few-shot examples).\n- **Benefit:** identical prompts in different requests skip embedding lookup.\n- **Implementation:** mark parts of prompt as \"cacheable\" at application level.\n\n**2. FlashAttention**\n- Reorder computation to maximize GPU utilization.\n- **Benefit:** faster attention with lower memory footprint.\n- **Compatibility:** works with KV cache — same keys/values, faster access.\n\n**3. KV Cache Offloading**\n- Move less-used cache entries to CPU RAM or disk.\n- **Benefit:** larger effective cache size, cheaper memory tier.\n- **Trade-off:** increased latency on cache misses.\n\n**4. Quantization**\n- Store K/V values in lower-precision (e.g., 8-bit instead of float16).\n- **Benefit:** 2–4× smaller cache memory.\n- **Consideration:** reduced accuracy, need fine-tuning.\n\n**5. Attention Sink / Alibi**\n- Special tokens (e.g., <sink>) stay in cache across many steps.\n- **Benefit:** maintains context for long-range dependencies.\n- **Use case:** long conversation handling, retrieval-augmented generation.\n\n**Golden rule:** caching is optimization, not magic. Profile your workload: where are the cache hits? Where are the misses? That's where to optimize next.\n"},{"id":"caching-summary","title":"Caching Summary","icon":"summarize","order":6,"file":"okf/llms/caching/caching-summary.md","body":"---\ntype: Section\ntitle: Caching Summary\ndescription: Caching - Caching Summary\ntags: [caching-summary,caching,llms]\ntimestamp: 2026-08-26T23:11:16.584Z\nsection: caching-summary\nguide: caching\nphase: llms\nicon: summarize\norder: 6\n---\n\n# Caching Summary\n\n**Icon:** summarize\n\nCaching transforms the O(n³) attention computation across tokens into O(n) per step — the single biggest performance win in modern LLMs. It enables everything from real-time chat to billion-parameter model serving.\n\n**The three caching pillars:**\n1. **KV Cache:** attention keys and values (the heavyweight, stable data).\n2. **Prompt Cache:** static embedding components (the repeatable input).\n3. **Model Cache:** intermediate representations (the reusable computations).\n\n**Key insights:**\n- Caching is always beneficial: it trades memory for compute.\n- The right cache strategy depends on latency, memory, and cost constraints.\n- Advanced techniques (quantization, offloading, attention variants) keep pushing the frontier.\n\n**Golden rule:** start simple (full KV cache). Optimize based on bottlenecks: sliding window for memory, quantization for cost, offloading for scale. Measure cache hit rates — they tell you exactly where optimization is needed.\n"}]}]}]}