🗺️ Roadmap: The Exam Prep & Certification¶
Below is the syllabus timeline of lessons in this module. Track your study progress here.
📈 Progress¶
📁 36 NCA-AIIO Blueprint Deep Dive — Domain-by-Domain Analysis¶
- 36.1a Domain 1 analysis: AI Infrastructure Fundamentals — coverage gaps and key concepts
- 36.1b Domain 2 analysis: GPU Architecture and Systems — exam-critical hardware specs
- 36.1c Domain 3 analysis: Networking for AI — InfiniBand vs Ethernet decision trees
- 36.1d Domain 4 analysis: Storage for AI — GPUDirect Storage scenario questions
- 36.1e Domain 5 analysis: Software Stack and Containerization — NGC and NVAIE
- 36.1f Domain 6 analysis: The Cluster Orchestration — GPU Operator and MIG deep dives
- 36.1g Domain 7 analysis: Monitoring and Operations — DCGM and Xid error scenarios
- 36.2a GPU architecture comparisons: when to use A100 vs H100 vs L40S
- 36.2b MIG partitioning: profile selection for given workload scenarios
- 36.2c NVLink vs PCIe: bandwidth calculations and topology selection
- 36.2d InfiniBand vs Ethernet: scenario-based selection questions
- 36.2e DCGM vs nvidia-smi: when each tool is appropriate
- 36.2f GPUDirect Storage: architecture questions and performance impact
📁 37 Comprehensive Practice Assessments¶
- 37.1a Part I–III: Infrastructure Foundations — 80 questions with detailed rationales
- 37.1b Part IV–V: GPU Architecture and Networking — 90 questions with rationales
- 37.1c Part VI–VII: Storage and Software Stack — 70 questions with rationales
- 37.1d Part VIII–IX: Orchestration and Operations — 80 questions with rationales
- 37.1e Parts X–XI: Cloud and Security — 40 questions with rationales
- 37.2a Mock Exam 1: Foundational focus — infrastructure, hardware, and networking
- 37.2b Mock Exam 2: Software and operations focus — containers, K8s, monitoring
- 37.2c Mock Exam 3: Mixed adaptive exam — full NCA-AIIO simulation with score analysis
- 37.3a Case Study: Designing a 128-GPU InfiniBand cluster for LLM pre-training
- 37.3b Case Study: Diagnosing a training job that suddenly drops GPU utilization to 20%
- 37.3c Case Study: Migrating a bare-metal AI cluster to Kubernetes with the GPU Operator
- 37.3d Case Study: Sizing storage for a 10 TB dataset training workload with GPUDirect Storage
📁 38 Reference Appendices and Quick-Reference Guides¶
- 38.1a GPU Specifications Comparison Table: V100 → A100 → H100 → H200 → B200
- 38.1b NVLink bandwidth progression table by generation
- 38.1c InfiniBand speed tiers: HDR/NDR bandwidth and port configuration table
- 38.2a nvidia-smi command cheat sheet: 30 most useful flags and queries
- 38.2b dcgmi cheat sheet: field IDs, health checks, and job monitoring commands
- 38.2c kubectl for GPU workloads: 25 essential commands for AI cluster operators
- 38.2d Linux administration cheat sheet: 40 commands every AI operator must know
- 38.3a A–F: AllReduce, Ampere, Baseboard Management Controller, BF16, CUDA, DPU, ECC, FP8...
- 38.3b G–N: GEMM, GPUDirect, HBM, HCA, Hopper, InfiniBand, MIG, MPS, NCCL, NVLink, NVSwitch...
- 38.3c O–Z: OAM, PFC, RoCEv2, RDMA, SM, SXM, Tensor Core, TensorRT, Triton, vGPU, Xid...
- 38.4a Official NVIDIA DLI (Deep Learning Institute) courses aligned to NCA-AIIO
- 38.4b Hands-on lab environment setup: building a local single-GPU practice environment
- 38.4c Exam registration: Pearson VUE scheduling, policies, and accommodation requests
- 38.4d Post-NCA-AIIO certification roadmap: NVIDIA Certified Professional (NCP) tracks