Computer Organization & Architecture

Unit 8: Introduction to Parallel Processing

From pipelining to multiprocessors โ€” master how modern CPUs execute billions of instructions per second through parallelism, pipelining hazards, and architectural innovations.

โฑ๏ธ 6 hrs theory + 4 hrs lab  |  ๐ŸŽฏ GATE ~2 marks  |  ๐Ÿ–ฅ๏ธ AWS 1 Crore Requests/sec

๐Ÿ’ผ Jobs this unlocks: CPU Design Engineer (โ‚น12โ€“25 LPA)  |  HPC Engineer (โ‚น10โ€“20 LPA)  |  VLSI Engineer (โ‚น8โ€“18 LPA)

Section A

Opening Hook โ€” 1 Crore Requests Per Second

๐Ÿข How Amazon AWS Handles 1 Crore Requests Every Second

When you click "Buy Now" on Amazon during the Great Indian Festival sale, your request is just one of 1 crore (10 million) requests processed every single second across AWS's global infrastructure. Behind this staggering scale lies the same principle you'll learn in this chapter โ€” parallel processing and pipelining.

Every modern CPU in AWS's data centres uses a pipelined architecture โ€” splitting instruction execution into stages so multiple instructions overlap like an assembly line. Their servers use multiprocessor systems with hundreds of cores executing tasks simultaneously. The Intel Xeon and AMD EPYC chips powering AWS use superscalar, out-of-order execution โ€” processing 4โ€“6 instructions per clock cycle.

What if YOU understood how this works? What if you could design pipelined processors, calculate speedups, and understand why your โ‚น50,000 laptop has 8 cores but your program only uses 1? That's exactly what this chapter teaches you.

๐Ÿ‡ฎ๐Ÿ‡ณ AWS India (Hyderabad)๐Ÿ‡ฎ๐Ÿ‡ณ Intel India (Bengaluru)๐Ÿ‡ฎ๐Ÿ‡ณ AMD India (Hyderabad)๐Ÿ‡ฎ๐Ÿ‡ณ Arm India (Bengaluru)๐Ÿ‡ฎ๐Ÿ‡ณ Qualcomm India๐Ÿ‡ฎ๐Ÿ‡ณ ISRO PARAM
India's PARAM Siddhi AI supercomputer at C-DAC Pune ranked 63rd in the world's Top500 supercomputers. It uses 42,000+ GPU cores running in parallel โ€” processing 5.27 petaflops (5.27 ร— 10ยนโต floating-point operations per second). That's the power of parallel processing!
Section B

Learning Outcomes โ€” Bloom's Taxonomy Mapped (12 Outcomes)

Bloom's LevelLearning Outcome
๐Ÿ”ต RememberList the 5 stages of a classic instruction pipeline (IF, ID, EX, MEM, WB) and define each stage's function
๐Ÿ”ต RememberState Flynn's four classifications (SISD, SIMD, MISD, MIMD) with one real-world example for each
๐ŸŸข UnderstandExplain pipeline hazards (structural, data, control) and how forwarding, stalling, and branch prediction resolve them
๐ŸŸข UnderstandDescribe the difference between shared-memory and distributed-memory multiprocessor organisations
๐ŸŸก ApplyCalculate pipeline speedup using S = nk/(k+nโˆ’1) and verify with worked numerical examples
๐ŸŸก ApplyApply Amdahl's Law to compute maximum speedup given fraction of parallelisable code and number of processors
๐ŸŸ  AnalyseDetect RAW, WAR, and WAW data hazards in a given instruction sequence and insert stalls/forwarding paths
๐ŸŸ  AnalyseCompare superscalar vs VLIW architectures on issue width, hardware complexity, and compiler dependency
๐Ÿ”ด EvaluateEvaluate the trade-offs between deeper pipelines (more stages) and increased hazard penalties in real processors
๐Ÿ”ด EvaluateAssess whether adding more processors is cost-effective using Amdahl's Law for a given workload
๐ŸŸฃ CreateDraw a complete space-time diagram for n instructions in a k-stage pipeline with hazard annotations
๐ŸŸฃ CreateDesign a parallel processing solution for a given real-world problem (e.g., image processing, web serving)