Volume II · Orientation
A chip begins as a decision system
Architecture is the negotiated answer to what work matters, what must move, and which constraints cannot be escaped.
AI workloads combine dense arithmetic, memory access, synchronization, communication, control, and storage. Training and inference weight these behaviors differently, while model shape, numerical format, batch size, latency target, reliability, and software ecosystem change the balance again. Hardware teams translate that workload envelope into throughput, memory, interconnect, power, area, and programmability targets.
PyTorch documents data, sharded-data, tensor, and pipeline parallelism as complementary distributed-training approaches. Those software choices are also hardware questions: they influence local memory capacity, accelerator-to-accelerator links, collective operations, host coordination, and how failure or stragglers affect the whole job.Claim
Design boundary: There is no universally best accelerator. A design is good only relative to workloads, deployment constraints, software readiness, cost, and time.
Turn a workload claim into a design contract
- Characterize— Describe operation mix, tensor shapes, precision, sparsity, locality, sequence behavior, and synchronization.
- Constrain— Set latency, throughput, power, thermal, memory, reliability, security, and cost boundaries.
- Model— Estimate compute, movement, storage, communication, and control bottlenecks across representative cases.
- Prototype— Compare architecture choices with traceable assumptions and sensitivity tests.
- Freeze— Record the benchmark, software stack, corner cases, and acceptance rule that authorize implementation.



