Skip to content
#

model-efficiency

Here are 12 public repositories matching this topic...

Token cost is a design problem, not a billing problem. Most LLM cost overruns come from architectural waste, not model pricing. This tool is a token waste profiler that helps you understand where your tokens are going and which ones are useless.

  • Updated Aug 9, 2026
  • Python

A brain-inspired language model that gets cheaper as it gets bigger. Top-1 spiking experts + an offline "sleep" phase that rewires the network → up to 7.1× less serving energy and ~1/43 the active compute of a dense model its size, while staying quality-competitive.

  • Updated Jul 31, 2026
  • Python

Improve this page

Add a description, image, and links to the model-efficiency topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the model-efficiency topic, visit your repo's landing page and select "manage topics."

Learn more