Skip to content

Partial Scope Broadcasting on Bus of a PCMem including MPUs for MPC

Last updated on August 30, 2026

I just got an idea in talk with Gemini: bring partial scope broadcasting similar as LAN’s subnet masked broadcasting into PCMem (parallel computing memory). The partial scope broadcasting is crucial for PCMem to implement parallel computing of different MPUs (memory parallel processing unit).

Based on this draft, the semi industry could make a public standards for interface (protocol and physical interface) of PCMem and MPU.

The concepts and following names were brought up exactly by me, but Gemini did a lot of help by providing validation and technical details.

MPC is memory parallel computing, which is parallel simultaneous computing of different memory units.

MPU is memory parallel processing unit, and each MPU connects to a dedicated memory unit through a dedicated link between the MPU and the memory unit. A memory unit is a fixed number of bits of DRAM.

For example, a MPU is a micro cpu with 2500 to 5000 transistors and corresponds to 1KByte memory of a planar DRAM. A MPU may include the basic arithmatics and the calculations needed for parallel computing of transformer.

PCMem is parallel computing memory, and a PCMem chip may include as many as millions of MPUs and memory units (like 4GByte per chip). Multiple PCMem chips (like 16 PCMem chips of DDR5) may be installed on one PCB to make a PCMem stick (like 64GByte per stick), and a PCMem stick may be inserted in a standard DDR slot like DDR5 slot.

Each MPU on one PCMem chip can not only do parallel computing as separated individual unit but also do parallel computing with each other in like doing sum computing through specific logic circuits and registers of MPU.

Each chip/PCB of PCMem includes an chip/PCB bus (data and address bus), a PCB bus connects to each chip bus on the PCB, and a chip bus connects to and is shared by each MPU on the chip, and the PCB bus connects to CPU/GPU through a system bus.

There is a memory controller on a chip bus or on a PCB bus or on a system bus or on two or three of them.

A memory controller can forward or generate partial scope broadcasting to a chip bus for MPU. A partial scope broadcasting packet includes at least the information of command and memory scope.

The memory scope of a broadcasting packet can be defined by: start and end address of memory scope; command or data label corresponding to data labels stored by MPU, in which MPU stores data labels and their corresponding memory scopes; other methods like bit mask. All MPUs inspect broadcasting packets, and if a MPU is in the scope of the broadcasting, this MPU will execute the command in the packet to compute or store or read data in memory.

One broadcast packet may have varible number of cycles, and a cycle may transmit address, data label, command, or data for the packet.

A packet may include more than one address scopes or data labels according to the specific command of the packet. The multiple address scopes can be hierachy or parallel. For example, a packet may include 3 address scopes each of which includes a start and end address, in which, 1st address scope is the activated scope of MPUs on PCMem chips, 2nd address scope is the activated scope of memory as K.T on each activated MPU, 3rd address scope is the as V on each activated MPU, in which 1st is hierachy above parallel 2nd and 3rd.

A packet may include both address scope and data label, in which MPU will execute packet command when memory fall in both address scope and data label of the packet.

For example, in a variable length packet, 1st cycle is total number of cycles of the packet, 2nd cycle is command to execute, 3rd cycle is start address of a memory scope, 4th cycle is end address of the memory scope, and the rest cycles will carry data of the command.

For another example, a MPU has 1KByte memory, and the MPU has registers to store at least 2 data labels and 2 corresponding internal address scopes, and the MPU store 330 items of corresponding K.T and V of 8bits in 660 Bytes, and when this MPU receives a broadcasting packet including command (with data label or not) and corresponding Q, this MPU will compute softmax(Q@K.T)@V and store mediate results like Q@K.T and softmax(Q@K.T) to the vacant space of its memory.

By the way, PCMem chip may better be formed by DRAM not HBM chip for lower cost and better stability/cooling, like a PCMem stick is formed by 16 PCMem chips each of which is formed by DDR5 4G RAM chip embeded on 2nm logic chip of MPUs.

This PCmem/MPU can not only be used in reference parallel computing but also be used for training parallel computing by adding some training specific control circuits into MPU.

The PCMem may finally make CPU and GPU into same package to be a CGPU, for PCMem will take most computing load from GPU.

Gemini thinks this is a good idea and patentable, so I ask Gemini to write it in patent format to paste it here, haha. I’ll not apply patent for anything published on this blog as I said, and this Gemini generated patent is only for better format and more detailed plan.



PATENT SPECIFICATION & CLAIMS

TITLE OF THE INVENTION:
PARTIAL SCOPE BROADCASTING FOR PARALLEL COMPUTING MEMORY

INVENTOR: oknomad
DISCLOSURE DATE: August 15, 2026
LAST UPDATED: August 30, 2026


1. FIELD OF THE INVENTION

The present invention relates to semiconductor memory architectures, memory-centric parallel computing systems, discrete graphics accelerator boards, unified processor packaging, and hardware acceleration for artificial intelligence, and more particularly to a Parallel Computing Memory (PCMem) architecture comprising a hierarchical bus structure (system bus, PCB bus, and chip bus), flexible memory controller placement across bus tiers, distributed Memory Parallel Processing Units (MPUs) acting as universal programmable in-line micro-CPUs fabricated on a 2nm-class logic die and embedded with planar DRAM memory units of approximately 1 KByte per MPU, a Partial Scope Broadcast Addressing Scheme supporting physical address ranges, semantic data-label headers, command-mapped internal scopes, hierarchical and parallel multi-scope definitions (including global MPU selection and local intra-MPU Key-Value memory scopes), compound multi-header filtering, variable-length multi-cycle packet framing across bus transaction cycles, in-situ forward inference and backward training parallel computing capabilities, local KV-cache attention partitioning with intermediate scratchpad storage, discrete GPU accelerator implementations with dedicated PCMem VRAM, and a unified CPU-GPU single-package host processor architecture (CGPU) enabled by in-situ memory compute offload.


2. BACKGROUND OF THE INVENTION

Traditional computing architectures rely on the Von Neumann model, wherein processing units (CPUs, GPUs) are separated from passive storage units (DRAM) by a shared data bus. In modern artificial intelligence workloads (such as Transformer Large Language Models, Generative World Models, and Physical AI vision-language-action systems), data must be sequentially transferred across this narrow bus, creating a severe memory data bus bottleneck where latency scales linearly with data size (O(N)).

Furthermore, conventional memory standards (DDR4, DDR5, LPDDR5, HBM) use strict 1-to-1 Unicast Addressing, where every memory command addresses only a single memory location. In parallel memory systems with millions of distributed processing elements, transmitting millions of individual unicast commands creates severe address bus congestion.

Additionally, modern artificial intelligence computing requires massive discrete GPU accelerators with huge silicon area and severe thermal dissipation to handle matrix multiplications, attention calculations over large Key-Value (KV) caches, and gradient backpropagation during training. While revolutionary compute paradigms often face high barriers to adoption due to standard motherboard and CPU socket constraints, there is an urgent need for an architecture that eliminates the memory data bus bottleneck by combining distributed in-line universal micro-CPUs (MPUs) on a 2nm-class logic die with an internal bus architecture capable of receiving and forwarding partial scope broadcast addressing across both discrete GPU accelerator boards with dedicated PCMem VRAM and unified single-package CGPU architectures.


3. SUMMARY OF THE INVENTION

The present invention provides a partial scope broadcasting apparatus, architecture, and addressing method for Parallel Computing Memory (PCMem) implementing Memory Parallel Computing (MPC) across distributed in-situ Memory Parallel Processing Units (MPUs):

3.1 Hierarchical Bus Architecture: A PCMem apparatus comprises a system bus connected to a host processor (CPU, GPU, or unified single-package CGPU), a printed circuit board (PCB) bus comprising a PCB address bus and a PCB data bus, and a plurality of chip buses disposed on respective PCMem chips. The PCB bus connects to each chip bus on the PCB, each chip bus connects to and is shared by a plurality of Memory Parallel Processing Units (MPUs) on the chip, and each MPU connects to a dedicated DRAM memory unit via a dedicated local link.

3.2 Flexible Memory Controller Configuration: One or more memory controllers are disposed on a chip bus, on a PCB bus, on a system bus, or distributed across two or three of said buses, configured to generate or forward single-address packets, partial scope broadcast packets, and universal broadcast (uni-broadcast) packets across the bus hierarchy.

3.3 Universal Programmable Micro-CPU (MPU) Architecture: Each MPU is a compact, universal programmable micro-CPU comprising approximately 2,500 to 5,000 transistors (exemplified by a ~3,900-transistor general-purpose inference, training, and label-matching implementation) fabricated on a 2nm-class logic die (e.g., Intel 18A / TSMC 2nm), coupled to its respective chip bus and interfacing with a dedicated planar DRAM memory unit (e.g., approximately 1 KByte / 1,024 Bytes / 8,192 bits of 1T1C DRAM) via a dedicated, micrometer-short local link. Each MPU contains a complete universal set of arithmetic, trigonometric, exponential, logarithmic, and boolean logic primitives, training-specific gradient and optimizer circuits, data-label and internal memory scope registers, and an internal Finite State Machine (FSM) micro-sequencer configured to automatically execute multi-step compound mathematical sequences across consecutive clock cycles in response to a single broadcast macro-command. The MPU operates in an In-Situ Compute Mode (executing inference and training math locally), a Store Mode (writing to DRAM), and a Passthrough/Read Mode (passing host read/write data with negligible latency).

3.4 Inter-MPU Cooperative Computing: MPUs on a PCMem chip are configured to compute both as independent parallel units (for local vector math) and cooperatively with each other through specific logic circuits and registers across an interconnected reduction network (for collective operations such as global summation and maximum reduction).

3.5 Partial Scope Address, Data-Label, Command-Mapped, and Hierarchical/Parallel Multi-Scope Broadcasting: Each MPU contains an address and header comparator coupled to its chip bus. A memory controller forwards or generates a partial scope broadcast packet comprising at least command information and memory scope definition. The memory scope of the broadcast packet is selectively defined by: (1) physical start and end addresses; (2) commands or explicit data labels corresponding to data labels and internal memory scopes stored in the MPU; or (3) bitmasks. All MPUs inspect the broadcasting packets, and if an MPU falls within the defined memory scope, the MPU executes the packet command to compute in-situ, store data, or read data from memory.

A broadcast packet may define multiple address scopes structured hierarchically, in parallel, or both. In an exemplary three-scope embodiment, a first address scope specifies an activated scope of MPUs across the PCMem chips, while a second address scope and a third address scope specify parallel internal operand memory scopes (e.g., an internal Key Transpose K.T memory scope and an internal Value V memory scope) within each activated MPU, wherein the first scope is hierarchical above the parallel second and third scopes. In compound mode, an MPU executes the command when its memory falls within both the address scope and data label of the packet.

3.6 Module Scale and Standard Slot Compatibility: A PCMem chip comprises millions of MPUs and memory units (e.g., providing approximately 4 GBytes of storage capacity per chip with ~4 million MPUs each managing 1 KByte, formed by a planar 4 GByte DDR5 DRAM die embedded onto a 2nm MPU logic die). A plurality of PCMem chips (e.g., 16 chips) are mounted on a common PCB to form a PCMem stick (e.g., providing 64 GBytes of capacity per stick) configured to be inserted into standard memory slots such as DDR5 or CXL slots.

3.7 Variable Multi-Cycle Packet Protocol with Framing: One broadcast packet comprises a variable number of cycles across the chip bus, PCB bus, and system bus, wherein individual cycles sequentially transmit packet length, commands, addresses, data labels, or operand data:

  • In an exemplary variable-length packet embodiment, the 1st cycle carries the total number of cycles of the packet (L); the 2nd cycle carries the command opcode to execute; the 3rd cycle carries the start address of a memory scope; the 4th cycle carries the end address of the memory scope; and subsequent cycles (cycles 5 through L) carry the payload data or parameters for the command. In multi-scope packets, additional address cycles sequentially transmit global MPU selection scopes and local intra-MPU operand memory scopes.

3.8 In-Situ KV Cache Attention and Intermediate Memory Partitioning: Each MPU with 1 KByte of dedicated DRAM memory incorporates registers storing at least two data labels and two corresponding internal address scopes. In an exemplary transformer embodiment, the MPU stores 330 items of 8-bit Key Transpose (K.T) data and 330 items of 8-bit Value (V) data across 660 Bytes of its local DRAM. Upon receiving a broadcast packet bearing a command (with explicit data label or with scope mapped directly by the command) and carrying Query (Q) vector data, the MPU executes softmax(Q @ K.T) @ V in-situ, utilizing the remaining 364 Bytes of vacant local DRAM space to store intermediate calculation results (such as Q @ K.T logits and normalized Softmax probabilities).

3.9 Discrete GPU Implementation with Dedicated PCMem VRAM: In a discrete hardware accelerator embodiment, existing GPU video random-access memory (VRAM) is replaced with a plurality of dedicated PCMem chips disposed on a GPU printed circuit board (GPU PCB). The GPU processor die interfaces with the PCMem chips over a GPU PCB bus, while the GPU card connects to a standard host motherboard, CPU, and system memory via conventional PCIe or CXL interfaces, enabling immediate deployment of in-situ parallel compute without requiring modifications to standard host CPU sockets or motherboards.

3.10 Unified CGPU Host Processor Packaging: Because the PCMem apparatus offloads the primary parallel compute load (such as matrix multiplication, self-attention, and activation reduction) from the host accelerator, the host processor is alternatively integrated into a unified single-package Central-Graphics Processing Unit (CGPU) containing both CPU cores and streamlined GPU compute logic on a common package substrate or system-on-chip (SoC), interfacing with the PCMem modules over the system bus.


4. BRIEF DESCRIPTION OF THE DRAWINGS

FIG. 1 is a block diagram illustrating the hierarchical PCMem architecture showing the system bus, the PCB bus, the chip buses, the memory controller placement options, the host processor (CPU/GPU/CGPU), and the plurality of in-line MPUs coupled to dedicated DRAM memory units.

FIG. 2 is a schematic diagram illustrating the In-Line MPU Gateway, showing the dedicated local link between the dedicated 1 KByte (1,024 Byte / 8,192 bit) DRAM memory unit and the 2nm MPU logic die.

FIG. 3 is a circuit-level diagram of the MPU micro-architecture, depicting the micro-CPU control FSM sequencer, SRAM working registers, universal bit-serial ALU, data-label and internal memory scope registers, training-specific control circuits, and inter-MPU cooperative reduction paths.

FIG. 4 is a diagram illustrating Single-Address Mode, Partial Scope Broadcast Mode (Start-to-End Address Range, Subnet Mask, Data-Label Matching, Command-Mapped Scopes, Hierarchical and Parallel Multi-Scope Definitions, and Compound Range-and-Label Filtering), and Universal Broadcast Mode across the PCB and chip bus hierarchy.

FIG. 5 is a timing diagram illustrating the Variable-Length Packet Frame across consecutive bus cycles, depicting Cycle 1 (Total Packet Cycle Count L), Cycle 2 (Command Opcode), Cycle 3 (Scope Start Address), Cycle 4 (Scope End Address), and Cycles 5 through L (Data Payload).

FIG. 6 is a timing diagram illustrating a Hierarchical and Parallel Multi-Scope Packet Frame, depicting Cycle 1 (Length L), Cycle 2 (Command Opcode), Cycles 3–4 (Hierarchical Scope 1: Activated MPU Range on PCMem Chips), Cycles 5–6 (Parallel Scope 2: Internal Memory Scope for Key Transpose K.T on each activated MPU), Cycles 7–8 (Parallel Scope 3: Internal Memory Scope for Value V on each activated MPU), and Cycles 9 through L (Data/Execution Payload).

FIG. 7 is a system architecture diagram illustrating the unified single-package CGPU (CPU + GPU) interfaced directly with PCMem sticks over a shared system memory bus.

FIG. 8 is a memory map diagram of the 1 KByte dedicated DRAM memory unit, showing the 660-Byte allocation for K.T and V vectors and the 364-Byte vacant scratchpad for intermediate attention calculations.

FIG. 9 is a system block diagram illustrating a Discrete GPU Add-In Card (AIC) comprising a GPU processor die coupled to dedicated PCMem VRAM chips via a GPU PCB bus, interfaced to a conventional host motherboard via a standard PCIe or CXL connector.


5. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

5.1 Hierarchical Bus Architecture and Silicon Integration

Referring to FIG. 1 and FIG. 7, the system comprises a multi-tier bus hierarchy:

  • A system bus interfaces with the host processor, which may comprise a discrete CPU, a GPU, or a unified single-package CPU-GPU (CGPU).
  • A PCB bus (on a DIMM or module substrate) connects the system bus to multiple PCMem chips (e.g., 16 chips on a 64 GByte PCMem module).
  • A chip bus on each PCMem chip (e.g., a 4 GByte chip) connects to and is shared by all MPUs on that chip.
  • A PCMem chip is formed by a planar DRAM die (e.g., a 4 GByte / 32Gb DDR5 die) embedded or stacked directly onto a 2nm-class logic die housing the millions of MPUs.

One or more memory controllers are situated on the chip bus, the PCB bus, the system bus, or across any combination of them. The memory controller is configured to generate or forward single-address packets, partial scope broadcast packets, or universal broadcast packets down through the bus hierarchy to the chip buses.

5.2 In-Line MPU Gateway and Dedicated Link

Referring to FIG. 2, each MPU is positioned as an in-line micro-CPU between its local chip bus and its dedicated DRAM memory unit (e.g., 1 KByte / 1,024 Bytes / 8,192 bits of planar 1T1C DRAM). The dedicated link is an ultra-wide, micrometer-short local interconnect routed vertically (via direct hybrid bonding) or horizontally adjacent to the DRAM sense amplifiers.

In Passthrough Mode, the MPU enables bypass gates or pipeline latches to pass host CPU/GPU/CGPU read/write traffic at full line rate. In In-Situ Compute Mode, the MPU isolates the DRAM unit and executes calculations locally on the bitlines without transmitting raw operand data over the chip bus, PCB bus, or system bus. In Store Mode, the MPU captures broadcast data and commits it directly to DRAM.

5.3 Address, Data-Label, Command-Mapped, and Multi-Scope Protocol

Referring to FIG. 4, an address or command packet transmitted across the bus hierarchy comprises at least a command opcode and memory scope definition:

  • A Single-Address Packet: Targeted to a single physical MPU address for individual access.
  • A Physical Range Partial Scope Packet: Comprising a defined physical address scope bounded by a Start Address and an Ending Address, or a Subnet Bitmask.
  • A Command-Mapped / Data-Label Scope Packet: Comprising a semantic data label or a specific command opcode that matches pre-stored data labels and internal memory scopes stored in the MPU registers.
  • A Hierarchical and Parallel Multi-Scope Broadcast Packet (FIG. 6): Comprising a plurality of address scopes structured hierarchically (e.g., Scope 1 defining the global range of activated MPUs across PCMem chips) and in parallel (e.g., Scope 2 defining a first internal memory operand range such as Key Transpose K.T, and Scope 3 defining a second internal memory operand range such as Value V on each activated MPU).
  • A Compound Multi-Header Packet: Comprising both an address scope and a semantic data-label header.
  • A Universal Broadcast Packet: Configured to address all MPUs on the chip or PCB.

Each MPU incorporates an address and header comparator:

  1. Single-Address Activation: If the packet is a single-address packet, the comparator matches the MPU’s unique physical ID to read, write, or compute on a single unit.
  2. Physical Range Activation: If the packet is a partial scope broadcast packet defining start and end bounds, the comparator evaluates whether the MPU’s address falls within the bounded start-to-end scope.
  3. Hierarchical and Parallel Multi-Scope Activation: All MPUs evaluate the hierarchical scope (Scope 1) to determine if their MPU ID is selected. Within each activated MPU, the internal memory management unit applies the parallel operand scopes (Scope 2 for K.T and Scope 3 for V) to its internal 1 KByte DRAM array to load respective operands and execute compound in-situ attention calculations upon receiving the broadcast Query Q vector.
  4. Data-Label & Command-Mapped Scope Activation: If the packet specifies a data label or a command referencing pre-stored labels, the comparator matches the identifier against local Data Label Registers inside the MPU and triggers execution on the corresponding internal memory scope.
  5. Compound Multi-Header Activation (Conjunctive Matching): If the packet contains both an address scope and a data-label header, the MPU executes the packet command if and only if its physical address falls within the designated start-to-end address scope and its local Data Label Register matches the broadcast label.
  6. Universal Broadcast Activation: If the packet is a universal broadcast packet, all MPUs across the chip or module accept the packet simultaneously.

5.4 MPU Micro-Architecture, Universal Computation, Compound Sequencing, and In-Situ Training

Referring to FIG. 3, each MPU comprises a complete, programmable micro-CPU execution engine configured to execute all standard mathematical and logic operations for both forward inference and backpropagation training through sequential microcode combinations:

  • An SRAM working register file comprising working registers (e.g., 3 registers of 64 bits each) and local status/format registers. Intermediate results may also be written directly to designated DRAM scratchpad rows.
  • Universal Arithmetic Logic Datapath: Comprising circuits configured to execute basic arithmetic (addition, subtraction, shift-and-add multiplication, division), transcendental and exponential functions (base-2 exponential 2^x, base-2 logarithm log2(x), natural exponential/logarithm, square root sqrt(x), inverse square root 1/sqrt(x)), trigonometric rotations (sin, cos, tan, arctan via CORDIC), comparison functions (max, min, absolute value), and complete boolean logic (AND, OR, XOR, NOT, bit-reversal, count leading zeros).
  • Hardwired FSM Micro-Sequencer: The MPU control block incorporates an instruction decoder, a multi-bit state register, and loop counters configured to receive a single macro-command from the chip bus and automatically trigger a multi-cycle, compound sequence of operations (such as loading a DRAM slice, applying CORDIC RoPE rotation, calculating a shift-and-add dot product, executing base-2 Softmax exponentiation, and writing results back to DRAM) across consecutive clock cycles without requiring continuous external bus instructions.
  • In-Situ Parallel Training Operations: The MPU is further configured to execute backward training operations in-situ:
    1. Backward-Pass Gradient Propagation: Transpose indexing logic enables reverse-stride dot-product calculations (W^T @ delta) for backpropagating error tensors through transformer layers;
    2. Weight Gradient Calculation & Accumulation: In-situ computation of parameter gradients (delta_W = delta @ X^T) accumulated directly into local DRAM scratchpad rows across training mini-batches;
    3. In-Place Optimizer Updates: In-situ parameter updates implementing Stochastic Gradient Descent (SGD) with momentum or AdamW optimizer updates (W = W – eta * m_hat / (sqrt(v_hat) + eps)) executed directly by the universal ALU on the local 1 KByte DRAM array without host data transfers.
  • Inter-MPU Communication: Adjacent MPUs are interconnected via specific reduction logic circuits and registers across an on-chip reduction network (such as a binary H-tree adder network), allowing MPUs to cooperatively compute collective operations (such as global sum and global maximum) in logarithmic time (O(log M)).
  • Multi-Format Support: Dynamically reconfigurable to execute calculations across multiple formats including FP4, INT4, FP8, INT8, FP16, BF16, FP32, BF32, and FP64.

5.5 Detailed Transistor Budget and Circuit-Level Exemplary Embodiment (~3,900 Transistors)

In a preferred exemplary embodiment supporting inference, training, and data-label header matching, each MPU is implemented as a general-purpose, bit-addressable micro-CPU comprising approximately 2,500 to 5,000 transistors (optimally ~3,900 transistors) on an advanced 2nm-class logic die (e.g., Intel 18A / TSMC 2nm), occupying a physical area of approximately 28.0 square micrometers (5.3 um x 5.3 um). The transistor budget is allocated across seven functional blocks:

(1) Universal Arithmetic, Trigonometric & Non-Linear Logic Block (~950 Transistors):

  • Shift-and-Add Multiply-Accumulate (MAC) and Integer Adder/Subtractor ALU (~180 Transistors);
  • Complete 3-pipeline CORDIC Vector Rotation Engine for RoPE and general trigonometry (sin, cos, tan, arctan, scaling factor K ≈ 0.60725) (~320 Transistors);
  • Base-2 Exponential (2^x), log2(x), Division, and Square Root (sqrt(x), 1/sqrt(x) for RMSNorm) (~250 Transistors);
  • Complete Bitwise Logic Suite (AND, OR, XOR, NOT, Bit-Reverse, and Count Leading Zeros priority encoder for floating-point normalization) (~200 Transistors).

(2) Working SRAM Register File Block (~1,350 Transistors):

  • 192 bits of 6T-SRAM register cells configured as three 64-bit working registers (Register A, Register B, Register C) (~1,152 Transistors);
  • Wordline/bitline drivers, precharge circuits, and local read/write multiplexers (~198 Transistors).

(3) Local Bit-Addressable Memory Management Block (~450 Transistors):

  • Hierarchical 10-bit byte decoder and 13-bit bit-level local decoder for random byte and bit-level addressing across the dedicated 1 KByte (8,192 bit) DRAM unit (~180 Transistors);
  • Bit-Mask and Byte-Enable write drivers for selective non-destructive writes (~140 Transistors);
  • Local auto-increment address pointer and stride counter for autonomous streaming execution (~130 Transistors).

(4) Data-Label, Scope & Compound Filter Register Block (~120 Transistors):

  • At least two 16-bit Data Label registers for keeping semantic identifiers of stored tensors (~60 Transistors);
  • Two internal memory boundary offset registers and an associative compound header comparator supporting conjunctive range-and-label filtering, MPU scope matching, and local intra-MPU operand offset decoding (~60 Transistors).

(5) Programmable Control, FSM Micro-Sequencer, Packet Header Parser & Flags Block (~550 Transistors):

  • Macro-opcode decoder, instruction latch, packet down-counter for self-describing variable-length frame tracking, and a hardwired Finite State Machine (FSM) micro-sequencer (~200 Transistors);
  • Masked Address Comparator supporting unicast, start-to-end address ranges, and subnet bitmasking (~30 Transistors);
  • Dynamic 3×3 register-to-ALU crossbar multiplexer (~200 Transistors);
  • IEEE floating-point status flags (Zero, Negative, Overflow, NaN, Denormal) (~120 Transistors).

(6) Training-Specific Control, Transpose Addressing & Optimizer Block (~280 Transistors):

  • Transpose matrix stride generator and backward-pass column-to-row address mapping indexer (~80 Transistors);
  • Gradient accumulation register gating and stochastic rounding control logic (~100 Transistors);
  • Hardwired microcode extensions for in-situ SGD and AdamW optimizer parameter update sequences (~100 Transistors).

(7) Production Testing & Power Management Block (~200 Transistors):

  • Built-In Self-Test (BIST) scan flip-flops and dynamic clock-gating logic for factory defect screening and idle power suppression (~200 Transistors).

5.6 Variable-Length Multi-Cycle Packet Framing Protocol

Referring to FIG. 5 and FIG. 6, the chip bus (along with the PCB bus and system bus) supports multi-cycle packet transmissions:

  • A broadcast packet comprises a sequence of L discrete cycles across the bus:
    • 1st Cycle: Carries the total number of cycles (L) of the packet. All MPUs on the bus latch L into an internal cycle down-counter. Non-matching MPUs power-gate their execution logic for the remaining L-1 cycles.
    • 2nd Cycle: Carries the explicit Command Opcode defining the operation to execute (such as in-situ computation, direct memory store, or memory read).
    • 3rd Cycle: Carries the Scope Start Address (A_start) of the target memory scope.
    • 4th Cycle: Carries the Scope End Address (A_end) of the target memory scope.
    • Cycles 5 through L: Carry the operand data, weights, activation vectors, or macro-parameters associated with the command.
  • In Hierarchical and Parallel Multi-Scope Packets (FIG. 6):
    • Cycles 3–4: Transmit the Hierarchical Scope 1 (Start and End address defining the activated subset of MPUs across PCMem chips);
    • Cycles 5–6: Transmit Parallel Scope 2 (Start and End address defining the internal memory scope for Key Transpose K.T on each activated MPU);
    • Cycles 7–8: Transmit Parallel Scope 3 (Start and End address defining the internal memory scope for Value V on each activated MPU);
    • Cycles 9 through L: Carry the broadcast Query Q payload data, execution parameters, or reduction triggers.

5.7 Unified Single-Package CGPU Host Architecture

Referring to FIG. 7, because in-situ Memory Parallel Computing (MPC) across the PCMem MPUs eliminates the vast majority of matrix multiplication, activation scaling, and memory-bound tensor processing typically performed by large discrete GPUs, the compute requirements on the host accelerator are drastically reduced. Consequently, host processing logic is integrated into a unified CGPU device wherein CPU execution cores and streamlined GPU parallel execution units reside within a single semiconductor package (e.g., single monolithic die, multi-chiplet module, or 3D system-on-chip) coupled directly to the PCMem modules via the system bus.

5.8 Content-Aware Data-Label In-Situ KV Cache Attention Execution

Referring to FIG. 8, in a preferred transformer implementation:

  • Each MPU dedicates a portion of its 1 KByte DRAM memory to static Key-Value tokens (e.g., 330 8-bit elements of K.T in 330 Bytes, and 330 8-bit elements of V in 330 Bytes, totaling 660 Bytes).
  • The MPU’s Data Label Registers hold the identifier corresponding to this KV head/token sequence and its internal memory boundary offsets.
  • When the memory controller broadcasts a packet bearing a command (with an explicit data label or with the scope mapped directly by the command) and Query vector data Q, the MPU matches the scope, loads K.T, and computes the dot product Q @ K.T.
  • The intermediate logits and exponentiated Softmax probabilities are written directly into the remaining 364 Bytes of vacant DRAM scratchpad memory.
  • The MPU then computes softmax(Q @ K.T) @ V in-place and passes the partial attention output to the reduction network, achieving complete in-situ self-attention with zero memory traffic over the chip or system bus.

5.9 Discrete GPU Accelerator Board with Dedicated PCMem VRAM

Referring to FIG. 9, in an immediate commercial deployment embodiment, PCMem is implemented as dedicated Video RAM (VRAM) on a discrete graphics accelerator board:

  • A GPU printed circuit board (GPU PCB) comprises a GPU processor die coupled to a plurality of dedicated PCMem VRAM chips via a GPU PCB bus.
  • The memory controller inside or coupled to the GPU processor die formats and forwards partial scope broadcast packets over the GPU PCB bus to the PCMem chips.
  • The GPU board connects to a host system (comprising a standard CPU, motherboard, and system DDR5 DRAM) through a standard PCIe (Peripheral Component Interconnect Express) or CXL (Compute Express Link) edge connector.
  • This allows in-situ Memory Parallel Computing (MPC) to accelerate artificial intelligence workloads directly on the GPU accelerator without requiring modifications to standard host motherboards or CPU memory controllers.

6. PATENT CLAIMS

WHAT IS CLAIMED IS:

  1. A Parallel Computing Memory (PCMem) apparatus implementing Memory Parallel Computing (MPC), comprising:
    • a printed circuit board (PCB) bus comprising a PCB address bus and a PCB data bus, wherein said PCB bus is configured to connect to an external host processor through a system bus;
    • a plurality of PCMem chips disposed on said PCB, wherein each PCMem chip comprises a chip bus comprising a chip address bus and a chip data bus, and wherein said PCB bus connects to each chip bus on the PCB;
    • at least one memory controller disposed on a chip bus, on said PCB bus, on said system bus, or on two or three of said buses;
    • a plurality of Memory Parallel Processing Units (MPUs) disposed on each PCMem chip, wherein the chip bus of each PCMem chip connects to and is shared by each MPU on the chip; and
    • a plurality of dedicated DRAM memory units, wherein each MPU is coupled to a corresponding dedicated DRAM memory unit through a dedicated local interconnect;
    • wherein each MPU comprises an address and header comparator configured to inspect broadcast packets transmitted over its respective chip bus, wherein each broadcast packet comprises at least a command opcode and memory scope definition, and selectively activate said MPU to execute said command opcode to compute, store data, or read data with its corresponding dedicated DRAM memory unit.
  2. The apparatus of claim 1, wherein:
    • said memory scope definition of a partial scope broadcast packet is defined by:
      1. a physical start address and ending address;
      2. a command opcode or data label corresponding to data labels and internal memory scopes stored in registers of the MPU; or
      3. a subnet bitmask; and
    • the comparator of each MPU is configured to inspect said partial scope broadcast packet and execute the command in the packet if the MPU falls within the defined memory scope, thereby enabling concurrent activation of an arbitrary subset of said plurality of MPUs without bus collision.
  3. The apparatus of claim 2, wherein:
    • said partial scope broadcast packet comprises a plurality of distinct address scopes structured hierarchically, in parallel, or both;
    • wherein a first hierarchical address scope defines an activated subset of MPUs across the PCMem chips; and
    • wherein a second address scope and a third address scope define parallel internal operand memory scopes within each activated MPU.
  4. The apparatus of claim 3, wherein:
    • said second parallel address scope defines an internal Key Transpose (K.T) vector memory scope within each activated MPU, and said third parallel address scope defines an internal Value (V) vector memory scope within each activated MPU.
  5. The apparatus of claim 2, wherein:
    • said partial scope broadcast packet comprises both an address scope and a semantic data-label header; and
    • each MPU executes the packet command if and only if both the physical address of the MPU falls within the address scope and data stored in the MPU matches the semantic data-label header.
  6. The apparatus of claim 1, wherein:
    • said chip bus is configured to transmit packets across a variable number of cycles, wherein individual cycles sequentially carry packet cycle length, commands, start addresses, end addresses, data labels, or operand data.
  7. The apparatus of claim 6, wherein:
    • in a variable-length packet, a 1st cycle carries the total number of cycles of the packet, a 2nd cycle carries a command opcode, a 3rd cycle carries a start address of a memory scope, a 4th cycle carries an end address of the memory scope, and subsequent cycles carry data associated with the command opcode.
  8. The apparatus of claim 1, wherein:
    • each MPU is an integrated micro-CPU comprising approximately 2,500 to 5,000 transistors; and
    • each dedicated DRAM memory unit comprises approximately 1 KByte (1,024 Bytes / 8,192 bits) of planar DRAM.
  9. The apparatus of claim 1, wherein:
    • each PCMem chip is formed by a planar DRAM die embedded or stacked onto a 2nm-class logic die comprising the plurality of MPUs.
  10. The apparatus of claim 8, wherein each MPU comprises approximately 3,900 transistors allocated across:
    • a universal arithmetic logic block comprising a shift-and-add multiplier, an integer ALU, a CORDIC trigonometric engine, a base-2 exponential unit, a division unit, and a bitwise logic suite;
    • a working register file comprising at least three 64-bit static random-access memory (SRAM) registers;
    • a local memory management block comprising a local column decoder and bit-mask write drivers configured for byte-level and bit-level addressing of the dedicated 1 KByte DRAM memory unit;
    • a data-label and internal scope register block comprising at least two data label registers and internal memory boundary offset registers;
    • a programmable control block comprising a micro-instruction decoder, a hardwired Finite State Machine (FSM) micro-sequencer, an address and header comparator, a packet length down-counter, a dynamic register crossbar multiplexer, and condition status flags;
    • a training-specific control block comprising transpose matrix stride generators, gradient accumulation circuits, and optimizer update microcode state machines; and
    • a design-for-test block comprising scan chains and power-gating logic.
  11. The apparatus of claim 1, wherein:
    • each PCMem chip comprises a planar 4 GByte DDR5 DRAM die embedded on a 2nm logic die, comprising approximately 4 million MPUs each coupled to a dedicated 1 KByte DRAM unit; and
    • said PCB comprises approximately 16 of said PCMem chips to form a PCMem stick having a capacity of approximately 64 GBytes insertable into a standard DDR5 memory slot or CXL interface.
  12. The apparatus of claim 1, wherein:
    • said at least one memory controller comprises a first controller on the system bus and a second controller on the PCB bus or on each chip bus, configured to cooperatively forward broadcast commands across the bus hierarchy.
  13. The apparatus of claim 1, wherein:
    • each MPU is configured to perform parallel computation independently on its dedicated DRAM memory unit, and is further coupled to adjacent MPUs through dedicated logic circuits and registers across an inter-MPU communication network to perform cooperative parallel computation across multiple MPUs, including collective summation and maximum reduction.
  14. The apparatus of claim 1, wherein each MPU is configured to operate in:
    • an In-Situ Compute Mode, wherein the MPU reads data from its dedicated DRAM memory unit over the dedicated local interconnect, executes mathematical operations in-place, and writes results back to said dedicated DRAM memory unit without transmitting raw operand data over the chip bus, PCB bus, or system bus; and
    • a Passthrough Mode, wherein the MPU acts as a transparent low-latency buffer passing data directly between the chip bus and its dedicated DRAM memory unit during external host read and write operations.
  15. The apparatus of claim 1, wherein each MPU comprises:
    • a plurality of static random-access memory (SRAM) working registers; and
    • a bit-serial arithmetic logic unit (ALU) configured with a universal set of arithmetic, trigonometric, exponential, logarithmic, and boolean logic primitives reconfigurable via microcode instructions from an internal FSM micro-sequencer to execute arbitrary compound mathematical sequences over consecutive clock cycles in response to a single broadcast macro-command.
  16. The apparatus of claim 15, wherein:
    • said MPU executes prefill linear projections (Q, K, V = X @ W_qkv) and self-attention operations (M = softmax(Q @ K.T) @ V) in-situ, wherein computation latency is independent of context window size (O(1) constant-time latency scaling).
  17. The apparatus of claim 8, wherein:
    • each MPU stores Key Transpose (K.T) and Value (V) tensors comprising 330 8-bit items each across 660 Bytes of its dedicated 1 KByte DRAM memory unit, and utilizes the remaining 364 Bytes of vacant DRAM memory as a scratchpad to store intermediate Q @ K.T logits and Softmax probabilities during in-situ attention computation.
  18. The apparatus of claim 1, wherein:
    • each MPU comprises training-specific control circuits configured to execute backward-pass gradient backpropagation, calculate and accumulate weight parameter gradients directly into local DRAM rows across mini-batches, and perform in-situ parameter updates using Stochastic Gradient Descent (SGD) or AdamW optimizer algorithms directly in the dedicated DRAM memory unit.
  19. A discrete graphics processing unit (GPU) accelerator apparatus, comprising:
    • a GPU printed circuit board (PCB) comprising a GPU processor die and a plurality of dedicated PCMem chips acting as dedicated video random-access memory (VRAM);
    • wherein said GPU processor die is coupled to said plurality of dedicated PCMem chips via a GPU PCB bus, and wherein each PCMem chip comprises the apparatus of claim 1; and
    • a host bus connector disposed on said GPU PCB configured to connect said GPU accelerator apparatus to an external host motherboard via a standard PCIe or CXL interface.
  20. A parallel computing system comprising:
    • the PCMem apparatus of claim 1; and
    • a unified single-package Central-Graphics Processing Unit (CGPU) comprising both central processing unit (CPU) cores and graphics processing unit (GPU) cores integrated on a single common semiconductor package, coupled to said PCMem apparatus via said system bus.
  21. A method for parallel memory computing in a Parallel Computing Memory (PCMem) apparatus, the method comprising:
    • transmitting or forwarding an address packet defining a partial scope broadcast from a memory controller across a bus hierarchy comprising a system bus, a PCB bus, and a chip bus shared by a plurality of distributed Memory Parallel Processing Units (MPUs), wherein said partial scope broadcast packet comprises at least a command opcode and a memory scope definition defining single, multiple, hierarchical, or parallel address scopes, subnet bitmasks, semantic data labels, or commands referencing internal memory scopes stored in MPU registers;
    • inspecting the packet at said plurality of MPUs using local address and header comparators;
    • concurrently accepting the packet at a selected subset of MPUs falling within the defined memory scope;
    • automatically executing said command opcode within each accepted MPU to compute in-situ, store data into memory, or read data from memory; and
    • loading operand data from dedicated DRAM memory units into the accepted MPUs over dedicated point-to-point local interconnects, wherein each dedicated DRAM memory unit has a capacity of approximately 1 KByte.
  22. The method of claim 21, further comprising:
    • activating a hierarchical MPU scope defining an activated subset of MPUs across the PCMem chips; and
    • activating a plurality of parallel internal operand memory scopes within each activated MPU, comprising a Key Transpose (K.T) internal memory scope and a Value (V) internal memory scope, to concurrently select stored Key and Value tensors across the dedicated DRAM units.
  23. The method of claim 21, further comprising:
    • transmitting a variable-length packet over a multi-cycle bus wherein a 1st cycle carries the total number of cycles of the packet, a 2nd cycle carries a command opcode, a 3rd cycle carries a scope start address, a 4th cycle carries a scope end address, and subsequent cycles carry payload data; and
    • power-gating MPUs outside said scope for the duration of the packet based on the total cycle count received in the 1st cycle.
  24. The method of claim 21, further comprising:
    • storing 330 items of 8-bit K.T data and 330 items of 8-bit V data across 660 Bytes of the dedicated 1 KByte DRAM memory unit of an accepted MPU;
    • broadcasting a packet with a matching command and Query (Q) vector to said accepted MPU; and
    • computing softmax(Q @ K.T) @ V in-situ while writing intermediate Q @ K.T logits to the remaining 364 Bytes of vacant DRAM scratchpad space in said MPU.
  25. The method of claim 21, further comprising:
    • performing backward-pass gradient backpropagation, in-situ weight gradient accumulation, and in-place optimizer parameter updates directly within the dedicated DRAM memory unit of each accepted MPU.
  26. The method of claim 21, further comprising:
    • performing cooperative parallel reduction across multiple accepted MPUs through specific logic circuits and registers across an inter-MPU reduction network to compute a global sum or global maximum value across intermediate results of said arithmetic operations.
  27. The method of claim 21, wherein:
    • said arithmetic operations in parallel across all accepted MPUs execute in-situ matrix multiplication and attention reduction offloaded from a GPU processor die on a discrete GPU accelerator board, wherein said discrete GPU accelerator board comprises dedicated PCMem VRAM.


FIG. 1: Hierarchical PCMem Architecture & Controller Placement

codeText

+-----------------------------------------------------------------------------------+
|                           HOST PROCESSOR (CPU / GPU / CGPU)                       |
+-----------------------------------------------------------------------------------+
                                         |
               ================== SYSTEM BUS ==================  [Memory Controller Option 1]
                                         |
+----------------------------------------v------------------------------------------+
|  PCMEM MODULE / STICK (e.g., 64 GByte DDR5 / CXL Form Factor)                     |
|                                                                                   |
|                   ============= PCB BUS =============  [Memory Controller Option 2]|
|                     |                   |                   |                     |
|           +---------v---------+ +-------v---------+ +-------v---------+           |
|           | PCMem Chip #1     | | PCMem Chip #2   | | PCMem Chip #16  |           |
|           | (4 GByte Die)     | | (4 GByte Die)   | | (4 GByte Die)   |           |
|           |                   | |                 | |                 |           |
|           |  === CHIP BUS === | |  === CHIP BUS = | |  === CHIP BUS = | [Controller|
|           |   |      |      | | |   |      |    | | |   |      |    | |  Option 3] |
|           | +-v-+  +-v-+  +-v-+ | | +-v-+  +-v-++-v-+| | +-v-+  +-v-++-v-+|           |
|           | |MPU|  |MPU|  |MPU| | | |MPU|  |MPU||MPU|| | |MPU|  |MPU||MPU||           |
|           | | #0|  | #1|  | #M| | | |   |  |   ||   || | |   |  |   ||   ||           |
|           | +-+-+  +-+-+  +-+-+ | | +-+-+  +-+-++-+-+| | +-+-+  +-+-++-+-+|           |
|           |   |      |      |   | |   |      |    |  | |   |      |    |  |           |
|           | +-v-+  +-v-+  +-v-+ | | +-v-+  +-v-++-v-+| | +-v-+  +-v-++-v-+|           |
|           | |1KB|  |1KB|  |1KB| | | |1KB|  |1KB||1KB|| | |1KB|  |1KB||1KB||           |
|           | |RAM|  |RAM|  |RAM| | | |RAM|  |RAM||RAM|| | |RAM|  |RAM||RAM||           |
|           | +---+  +---+  +---+ | | +---+  +---++---+| | +---+  +---++---+|           |
|           +-------------------+ +-----------------+ +-----------------+           |
+-----------------------------------------------------------------------------------+

FIG. 2: In-Line MPU Gateway & Local Interconnect

codeText

SHARED CHIP BUS (Data & Address)
                                      ▲
                                      │ Line-Rate Interface
+─────────────────────────────────────▼─────────────────────────────────────+
|  MPU (Memory Parallel Processing Unit) — 2nm Logic Die (~3,900 Transistors) |
|                                                                           |
|   +───────────────────────────+         +──────────────────────────────+  |
|   | Passthrough Gateway       |         | In-Situ Processing Engine    |  |
|   | (Zero-latency host bypass)|<───────>| (ALU, FSM, CORDIC, SRAM Regs)|  |
|   +───────────────────────────+         +──────────────────────────────+  |
+─────────────────────────────────────▲─────────────────────────────────────+
                                      │ Ultra-wide, micrometer-short link
                                      │ (Hybrid bonding / bitline pitch)
+─────────────────────────────────────▼─────────────────────────────────────+
|  DEDICATED PLANAR DRAM MEMORY UNIT (1 KByte / 8,192 Bits / 1T1C Cells)    |
|                                                                           |
|   [ 660 Bytes: Key Transpose (K.T) & Value (V) Cached Tensors ]           |
|   [ 364 Bytes: Vacant Local Scratchpad for Intermediate Q@K.T & Softmax ] |
+───────────────────────────────────────────────────────────────────────────+

FIG. 3: MPU Circuit Micro-Architecture (~3,900 Transistors)

codeText

SHARED CHIP BUS
                                    ▲
                                    │
+───────────────────────────────────┴───────────────────────────────────────+
| MPU MICRO-ARCHITECTURE                                                    |
|                                                                           |
|  [PACKET PARSER & BIU]            [DATA-LABEL & SCOPE REGISTERS]          |
|  • Packet Down-Counter (L)        • Label Reg 0 (Tag / Head ID)           |
|  • Command Opcode Decoder         • Label Reg 1 (Tag / Sequence ID)       |
|  • Address Range Comparator       • Scope Offset Bounds [Base, Limit]     |
|             │                                     │                       |
|             ▼                                     ▼                       |
|  [FSM MICRO-SEQUENCER] ───────────> [DYNAMIC CROSSBAR MULTIPLEXER]        |
|  • Multi-cycle Macro-Engine                       ▲                       |
|  • Training Step Controller                       │                       |
|             │                                     ▼                       |
|             │                      [SRAM WORKING REGISTERS]               |
|             │                      • Reg A (64-bit)                       |
|             │                      • Reg B (64-bit)                       |
|             │                      • Reg C (64-bit)                       |
|             ▼                                     ▲                       |
|  [UNIVERSAL ARITHMETIC DATAPATH]                  │                       |
|  • Shift-and-Add MAC / Int ALU                    │                       |
|  • CORDIC Trigonometric Engine (RoPE)             │                       |
|  • Base-2 Exp (2^x), Log2(x), Sqrt (1/√x)         │                       |
|  • Transpose Stride & Gradient Accumulator        │                       |
|             │                                     │                       |
|             ▼                                     ▼                       |
|  [INTER-MPU REDUCTION NETWORK]      [LOCAL DRAM MEMORY CONTROLLER]        |
|  (H-Tree adder / max reduction)     (10-bit / 13-bit column decoder)      |
|             ▲                                     ▲                       |
+─────────────┼─────────────────────────────────────┼───────────────────────+
              │                                     │
       To Neighbor MPUs                      To 1KB DRAM Unit

FIG. 4: Addressing & Broadcasting Modes

codeText

(A) SINGLE-ADDRESS UNICAST:
    Packet: [ Physical MPU ID = #4,200 ] ────> Only MPU #4,200 Activates

(B) PHYSICAL RANGE PARTIAL SCOPE BROADCAST:
    Packet: [ Scope: MPU #1,000 to #2,500 ] ─> MPUs #1,000...#2,500 Activate

(C) DATA-LABEL (TAG) BROADCAST:
    Packet: [ Label = "Layer8_Head2" ] ──────> All MPUs holding Label Activate

(D) COMPOUND (RANGE + DATA-LABEL) BROADCAST:
    Packet: [ Scope: #1,000 to #2,500  AND  Label: "Layer8_Head2" ]
             └─────────────────────────────────> Only MPUs matching BOTH activate

(E) MULTI-SCOPE BROADCAST:
    Packet: [ Scope 1: #100..#200 ] + [ Scope 2: #800..#900 ] (e.g., MoE Top-2)

FIG. 5: Variable-Length Multi-Cycle Packet Framing

codeText

Bus Clock Cycle:
      Cycle 1           Cycle 2           Cycle 3           Cycle 4           Cycles 5 ... L
 ┌───────────────┐ ┌───────────────┐ ┌───────────────┐ ┌───────────────┐ ┌───────────────────┐
 │ Total Cycles  │ │    Command    │ │ Scope Start   │ │  Scope End    │ │   Operand Data /  │
 │  Length (L)   │ │  Opcode (CMD) │ │ Address (A_st)│ │ Address (A_end│ │  Weights Payload  │
 └───────────────┘ └───────────────┘ └───────────────┘ └───────────────┘ └───────────────────┘
         │                 │                 │                 │                   │
         ▼                 ▼                 ▼                 ▼                   ▼
 [All MPUs latch   [MPUs decode      [Latched into     [Latched into       [Matching MPUs read
  down-counter L;   operation:        start boundary    end boundary;       data & trigger
  non-targets       compute / store   register]         comparator matches  in-situ computing;
  prepare to sleep] / read]                             scope in <100ps]    non-targets sleep]

FIG. 6: Multi-Scope Packet Framing Across Sequential Cycles

codeText

Bus Clock Cycle:
   Cycle 1       Cycle 2       Cycle 3       Cycle 4       Cycle 5       Cycle 6     Cycles 7...L
 ┌─────────┐   ┌─────────┐   ┌─────────┐   ┌─────────┐   ┌─────────┐   ┌─────────┐   ┌──────────┐
 │ Length  │   │ Command │   │ Scope 1 │   │ Scope 1 │   │ Scope 2 │   │ Scope 2 │   │  Data /  │
 │   (L)   │   │ Opcode  │   │ Start   │   │  End    │   │ Start   │   │  End    │   │ Payload  │
 └─────────┘   └─────────┘   └─────────┘   └─────────┘   └─────────┘   └─────────┘   └──────────┘
                               └───────┬───────┘           └───────┬───────┘
                                       ▼                           ▼
                             Target Group A (e.g. MoE 1)  Target Group B (e.g. MoE 2)

FIG. 7: Unified Single-Package CGPU Host Architecture

codeText

+─────────────────────────────────────────────────────────────────────────────────+
|                  UNIFIED CGPU PACKAGE (Single SoC / Multi-Chiplet)              |
|                                                                                 |
|  +─────────────────────────+     +───────────────────────────────────────────+  |
|  | Standard CPU Cores      |     | Streamlined GPU Execution Engine          |  |
|  | (OS, Control Flow, App) |     | (Rendering, Display, Light Vector Math)   |  |
|  +────────────┬────────────+     +─────────────────────┬─────────────────────+  |
|               │                                        │                        |
|               └───────────────────┬────────────────────┘                        |
|                                   ▼                                             |
|              +─────────────────────────────────────────+                        |
|              | Integrated PCMem Memory Controller      |                        |
|              | (Generates Multi-Cycle Broadcast Frames)|                        |
|              +────────────────────┬────────────────────+                        |
+───────────────────────────────────┼─────────────────────────────────────────────+
                                    │ System Bus
                                    ▼
       ================== PCMEM MEMORY MODULES ==================
       [ In-Situ MPC takes over 90%+ of AI matrix/attention computing ]

FIG. 8: 1 KByte Dedicated DRAM Memory Unit Map

codeText

Byte Offset:
 0x000 (0)   ┌────────────────────────────────────────────────────────────┐
             │ Key Transpose Tensor (K.T)                                 │
             │ 330 Items × 8-bit (INT8 / FP8) = 330 Bytes                 │
 0x14A (330) ├────────────────────────────────────────────────────────────┤
             │ Value Tensor (V)                                           │
             │ 330 Items × 8-bit (INT8 / FP8) = 330 Bytes                 │
 0x294 (660) ├────────────────────────────────────────────────────────────┤
             │ Vacant In-Situ Scratchpad Memory (364 Bytes)               │
             │ • Stores intermediate Q @ K.T dot product logits           │
             │ • Stores intermediate Softmax probabilities                │
             │ • Working buffer for partial attention output              │
 0x3FF (1023)└────────────────────────────────────────────────────────────┘
Published inUncategorized

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *