7 comments

  • kmike84 0 minutes ago
    This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)

    I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of he more optimized engines.

    3 main failure modes I observed in the engines:

    * Not using best available spec decoding * Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models) * Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline

  • sgtwompwomp 15 minutes ago
    This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
  • nateb2022 16 minutes ago
    Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
  • amirhesham 17 minutes ago
    Oh this is so cool. Curious about the business model, too.
  • kenzic 13 minutes ago
    How long does tuning take (on an M3 MacBook Pro for example)?
  • yolandac 9 minutes ago
    does it allow us to run larger models that weren't possible before?
  • p-e-w 23 minutes ago
    What is the business model?
    • anerli 1 minute ago
      We envision a future where workloads are hybrid. Average consumer hardware will be able to handle a lot with local models, but you’ll still want to use cloud models for harder tasks. Magnitude will make it seamless to switch between the two, even for the same tasks (without breaking your prefix cache). We’ll charge per token for our inference cloud, using the same efficiencies we unlock for local inference to pass the savings on to you.