Link everyday devices into one cluster and run models too big for any single machine.
Catalog snapshot
Clusters of 8–16 GB machines beat one box for 30B+ models.
Fetched 1
exo networks phones, Macs, and Linux boxes into one inference cluster and splits a model across them. It is the cheapest way to run models larger than any single machine you own. Expect tinkering, not a product.
device clustering · ChatGPT-compatible API · Apple Silicon · model sharding