← Topics

Local AI

— members
27 tweets
Columns:
# Tweet User Followers Views ▼ Ratio Engagement Posted
1
[video] Local AI is about to be competitive. I will do everything in my power to make it better than Claude desktop/claude code by end of year.
@0xSero ✓ 47.6K 213.0K 4.5x 4.2K May 1
2
[image] Best models for your hardware - 4gb to 12gb vram - VibeThinker-3B - smokes everything remotely close to its weight class. Challenging 30b models! Last version was also topping math benchmarks - 12gb to 24gb vram - Gemma-12B-coder Built on top of
@0xSero ✓ 53.6K 62.3K 1.2x 1.2K Jun 18
3
[image] Important: This is a summary of an amazing video by one of the best creators I know. Video in the first comment. If you've been struggling to setup a productive local environment I'll summarise, but you should watch the video. 1. Qwen3.6-27B with NO THINKING - 4bit - 16bit
@0xSero ✓ 47.9K 61.1K 1.3x 1.4K May 2
4
[image] Weekly best models for your hardware: ~~ 8 to 16gb ~~ Granite models are amazing: [NEW] - Gemma-E4B is a good general QA model - Qwen3.5-9B is the best at this level imo - ~~ 16 to 64gb ~~ Another
@0xSero ✓ 48.0K 56.6K 1.2x 1.3K May 2
5
[text] Breakthroughs: 1. Turboquant merged into vLLM 75% vram reduction for kvcache near losslsss 2. Someone merged M2.7 & M2.5 & got it to perform better than M2.7 3. 40% faster prefill on AMD strix halo (128gb for MoEs w 10B active params) 4. Megatrain 100B model trained on 1 GPU
@0xSero ✓ 43.7K 54.2K 1.2x 1.3K Apr 16
6
[image] Locally Part 1 - Apple Silicon Macs give you large pools of memory to run big models, but the token generation speed will be lower than most are used to. Macs are best with large MoEs that have low ACTIVE params. Basically when you see a model like Qwen3.5-397B-A17B this
@0xSero ✓ 44.8K 37.2K 0.8x 469 Apr 22
7
[video] To hell with Anthropic and their random policy changes. AI local is the future. This game was created using DeepSeek v4 entirely locally. The setup is not for the faint-hearted currently. DeepSeek v4 (came out a couple of days ago) GGUF 2-bit quant and llama.cpp patch from
@julianharris ✓ 5.6K 33.1K 5.9x 118 Apr 27
8
[video] local LLMs on Mac are getting good. but the workflow around models is still messy. you have models in: •⁠ ⁠LM Studio •⁠ ⁠Ollama •⁠ ⁠Hugging Face cache •⁠ ⁠random folders •⁠ ⁠manual downloads and every runtime expects something slightly different. we built
@sabeshbharathi ✓ 2.1K 24.8K 11.9x 220 May 24
9
[image] New build cooking - AMD threadripper 7965WX - Asus pro WS WRX90E-Sage SE - 256gb ddr5 - Corsair 9000D - 14 Noctua fans - 8tb NVMe I’m connecting an exhaust fan and pushing the air outside
@0xSero ✓ 57.6K 23.6K 0.4x 125 Jul 13
10
[text] Here’s what I’d recommend if you’re just getting started in AI, local or otherwise. 1. Work with the compute you have, even the dumbest LLMs can be useful if you treat them as a node in your system. Some basic problems of what could be useful to get you started - tag all
@0xSero ✓ 40.7K 23.1K 0.6x 753 Apr 3
11
[image] Cheapest competitive build on Nvidia 2 sparks = 8000$ Total specs - 256gb - 8tb - 546gb/s memory bandwidth - tons of flops --- models: - Deepseek-v4-flash - MiMo-v2.5-flash fp4 - MiniMax-M2.7 - Qwen3.5-397b-reap Flaws: - Low mem bandwidth - You need 2 for best perf
@0xSero ✓ 50.2K 22.1K 0.4x 317 May 15
12
[image] For any homelab friends. Tailscale + mullvad here’s a repo to help if you’re on Mac
@0xSero ✓ 57.5K 19.6K 0.3x 138 Jul 11
13
[image] Best harnesses for local models: 1. Droid: - Very good performance, forces the models to behave, you can wire in all your local LLMs very easily w BYOK - Allows you to use your local models as orchestrators/subagents so you can benefit from Cloud as models as well - Practically
@0xSero ✓ 41.6K 19.5K 0.5x 478 Apr 4
14
[image] AI workflows shouldn’t need 3 frameworks and 800 lines of Python glue Nika is a local-first AI workflow language + runtime, built in Rust 🦀 One YAML file. 4 verbs: infer, exec, invoke, agent Run GGUF/llama.cpp locally or cloud models. MCP, sandboxed permissions, traces +
@ThibautMelen ✓ 1.8K 17.5K 9.5x 17 Aug 20
15
[image] One thing I've done this year is: - Download all my X data from settings/account - Download all my youtube, gmaps, gmail, google from takout google com - Download all my personal data from Claude, ChatGPT - Export a copy of every AI session on Cursor Claude Code, Codex, Droid,
@0xSero ✓ 45.2K 16.9K 0.4x 470 Apr 23
16
[image] Based on my knowledge of: - AI labs - hardware manufacturers - model size performance - training process - training economics If you want to buy hardware for local AI you need 500-750gb of 500gb/s+ memory bandwidth to make serious use of local hardware at frontier levels
@0xSero ✓ 50.2K 16.1K 0.3x 267 May 15
17
[image] I recommend omp for your local AI if you have 48gb or more of memory, the advisor feature really helps steer it. The way it breaks down work and spawns subagents for everything, or supports setting a vision model is really how it should be done. Works with all your subs too
@0xSero ✓ 65.7K 15.9K 0.2x 322 Sep 1
18
[image] The smallest full-scale GLM-5.2 build on MLX has come out. You can download it now, and it’s also so good. It uses 2.32 bits per weight (219 GB) and fits on a 256 GB Mac with room left for context. Models this far down are usually barely usable. An anchor-guarded clip
@Alisvolatprop12 ✓ 61.8K 15.3K 0.2x 67 Jul 12
19
[image] To my fellow homelab friends, I need some help. My room is starting to cook, so I need to setup air conditioning (I already have a portable one but need something more pro) Essentially the rooms humidity is quite low and given I have cats there’s a lot of dust I need to clean
@0xSero ✓ 51.4K 14.9K 0.3x 200 May 29
20
[image] Guide to running BIG B0Is on your small hardware. 1. Use REAPs: up to 50% savings 2. Use quantisations: 75% savings - AWQ / GPTQ / W4A16 / FP8 = FAST inference - GGUF / EXL3 = Slow but just works - MLX = Best for apple 3. Use 8bit KV cache: 50-75% savings
@0xSero ✓ 43.4K 14.6K 0.3x 353 Apr 14
21
[image] Exl3 is the best quantization schema for local AI rn, especially if you have Blackwell lite (DGX/5090/4000/5000/6000) It was integrated into sparkinfer (fork of vLLM) and it’s running at 3x the speeds it was on exllamav3 It also works well when you have odd number of cards
@0xSero ✓ 61.3K 14.1K 0.2x 140 Aug 8
22
[image] This is the third version of KIMI-K3-alis-mlx. Compared to the previous V2, the weight size has been reduced by 211GB (from 948.5GB to 737.1GB), and the quality has been improved to be closer to the original through the DWQ technique. It still cannot run on a single M3 ultra Mac
@Alisvolatprop12 ✓ 62.0K 13.4K 0.2x 75 Aug 4
23
[image] Top 5 builds for AI inference in 2025-2026 I have spent around 12 months researching, building, experimenting, and bench-marking AI models, tools, hardware and costs. Top 3 picks will be the safest, best cost to performance ratios. The last 2 will be more interesting
@0xSero ✓ 48.0K 12.0K 0.2x 195 May 2
24
[image] I got 2 intel bad boys on Friday. At this point I’m struggling to find more power for all this. - 4x 6000s - 1x DGX Spark - AMD Strix - 4x 3090 - 2x intel arc b70 - Mac mini 16gb - MacBook Pro 32gb 544gb VRAM 300gb mixed Total = 844gb AI + 512gb ddr4 I need some guru help
@0xSero ✓ 51.8K 11.5K 0.2x 200 Jun 2
25
[text] We should get localmaxxing gang in here. I love what they're doing, going to start competing on it next week <3 Also does anyone here know how to make X bots the legal TOS way?
@0xSero ✓ 57.1K 10.0K 0.2x 143 Jul 5
26
[image] Why I built vllm-studio - Storing configurations for models and engines - Fully integrated agent with Pi - Model downloading, exploration, and config - Creating benchmark datasets Backend can be deployed to multiple servers, and the frontend can connect to all from anywhere.
@0xSero ✓ 50.9K 7.0K 0.1x 196 May 22
27
[text] Local AI is a human right, our children, families, neighbours, friends, and fellow humans deserve privacy and freedom.
@0xSero ✓ 42.2K 6.2K 0.1x 202 Apr 7
28
[image] Hi frens. My fren and his team have been seeking more providers for local compute. People can earn from- Any advice or help for setting up is available!
@TRACaveMan ✓ 378 743 2.0x 27 Sep 13
29
[image] Made a lot of progress on local AI on cheap hardware.. 20 tok/s Qwen 3.8 27b on a $100 P100.
@iamMrDuncan ✓ 920 552 0.6x 8 Sep 27
30
[image] Hey, check out this project—take a look at what the devs at @Parad0x_Labs are building. It’s not a chatbot. It’s a local agent runtime—the model proposes, the kernel executes and signs the receipt. Beta for macOS is out now. VOOL (@Parad0x_Labs) isn't a chatbot. It’s a local
@Degen922 ✓ 228 508 2.2x 14 Sep 20
31
[image] That was one law. Here is the finished constitution.
@NathanWalesVA ✓ 85 490 5.8x 4 Sep 26
32
[text] For true local AI cowboys: support the team!
@TRACaveMan ✓ 412 259 0.6x 3 Sep 29