Silicon
Compute, designed from the workload down.
Architecture for intelligence workloads — compute, memory and interconnect arranged around what models actually do, designed together with the compiler and runtime that drive them.
- 00Unresolved
- 01I/O ring
- 02Compute arrays
- 03Memory beside compute
- 04Interconnect
- 05A workload, placed
Abstract floorplan · not a Somber design · not to scale
Why silicon
A model's behaviour, its cost and its reach are usually solved by different people at different times. Somber treats them as one problem: the same lab works on the weights, the software around them and the compute beneath them, so each can constrain the others.
- 01
Model
Behaviour is decided in training and post-training.
- 02
Software
Latency, cost and reliability are decided by the software around it.
- 03
Compute
Throughput is decided by how work and data are scheduled.
- 04
Hardware
The limits of all three are set by physical constraint.
Compute principles
- 01
Movement costs more than arithmetic
Moving data dominates energy and time. Architecture starts with where data lives.
- 02
Locality is designed, not found
Memory sits beside the compute that needs it; placement is a first-class decision.
- 03
The compiler is part of the chip
Hardware and its software stack are designed together, or the hardware is not finished.
- 04
Measure the workload, not the peak
Performance is reported on real workloads, with the configuration that produced it.
Performance disclosure
Somber's silicon work is architecture research. Programs appear here with their development stage; performance figures appear only with full disclosure.
Every performance figure on these pages will state:
- 01Benchmark
- 02Workload
- 03Batch size
- 04Precision
- 05Power mode
- 06Software version
- 07Date
- 08Comparison configuration
- 09Caveats