I currently run the top 8 or so per day from r/LocalLLama through my RSS reader. Actually going to reddit is eww, so I want it to be worthwhile. I usually get enough to keep roughly up to date just with title and text / graphs and only go if it looks really interesting.
With the upcoming demise of reddit RSS, I’m wondering if people have recommended (hopefully RSS friendly) places roughly equivalent. I guess I could filter HN which may be useful for many tech interests, but I don’t know what I don’t know, so I thought I’d pose the question to the crowd.
HF and GitHub. I guess soon I’ll be looking for somewhere besides HF, too.
Huggingface is probably the primary source after Chinese forums
from what i am hearing and i really really really dont want to create an scvount but X (teitter). ftom what i gear, its a good resource. people are sharing good things
Reddit is eventually going to die of enshitification. The best way to survive it is to make this alternative work and promote it regularly (in a non spammy way) on reddit.
Agreed, but we’re not there yet.
what are you looking for? I could start posting more
also not OP but I am interested in local models usable on low-end hardware, models for tasks like document summarizing and text rewording, fuzzy search of various types of local data, and search agents such that I can ask it a vague natural-language question and it runs 17 web searches and comes back with a summary of what it found.
I’m not the OP, but I’m trying to figure out if local models are always painfully slow or if I’m missing something obvious in my tuning.
Stable Diffusion can whip up a picture in less time on the same hardware, than lama.cpp takes to decide to call an MCP function.
It seems like I must be missing something in my lama.cpp setup, but none of the guides I’ve read have clued me in to what I’ve done wrong.
Ollama performs similarly poorly on the same harsware, so I’ve probably managed to make the se mistake(s) at least twice.
Anyway, that’s the main thing I’m reading along for. Trying to increase my understanding until I catch my own mistakes.
I have a guide for performance tuning that should be a pretty good start
https://lemmus.org/post/24235317
Let me know if you have questions, or maybe just make a post asking how to optimize for your hardware and I’ll try to answer
I would suggest you don’t use Ollama https://sleepingrobots.com/dreams/stop-using-ollama/ If you want a GUI, Unsloth Studio is probably best and open source. LM Studio is good too but closed source.
I just use llama.cpp llama-server with the built-in Web UI
Also check the llama.cpp docs
https://github.com/ggml-org/llama.cpp/blob/master/tools/cli/README.md
I will study these. Thank you!
I do something similar to you, except I use redlib as a backend and a local qwen3.8:27b as an aggregator to build me a daily digest every day.
If we innovate a little we can survive the death of rss
I have my hermes agent run xvnc with chrome & playwright to navigate old.reddit.com with an old throwaway account. It has created a little script to reduce the DOM to an LLM-friendly size.
It also digests the consensus of the comments in the top posts.
Now there’s an interesting idea. Ooh, redlib even has a quadlet definition out of the box, don’t see that every day. I may well give that a spin, probably tack a rss feed on the end as well (creature of habit). Thanks.
if you actually intend to replicate my setup: [qwen3.8:27:nvfp4 --> vllm (concurrency efficiency)] --> [open-webui (api proxy)] --> [hermes-agent --> hermes-webui]
[…] denotes container boundaries.
The redlib instance also runs in a container on the same host. A skill teaches hermes-agent how to use redlib to open a subreddit and a separate instruction markdown file teaches it how to build the digest I want. a “cron” (hermes cron, not system cron) runs the job once a day, and the final step of the digest instructions is to email me the digest.
This was today’s, for example:
LocalLLaMA Daily Digest – 2026-10-04
Curated for a [REDACTED] LLM-backend sysadmin running vLLM + Open-WebUI on RTX 3090s and RTX PRO 6000 Blackwell. Focus: quantization (INT4/NVFP4/GGUF/EXL3/EXL2), inference serving, VRAM-limited local deployment, and reproducible configs.
Most frequently discussed topics
1. The rise of narrow “overfit” inference engines This was the loudest infrastructure thread today. A wave of deliberately non-general runtimes – Strata, Ninfer, DwarfStar, Splash, llamAmpere, gufo, Kyojin, TensorSharp – give up llama.cpp/vLLM’s generality to squeeze maximum throughput out of a handful of models on a single hardware family (often one GPU or the Strix Halo APU). For a vLLM/Open-WebUI operator, the question is whether it is worth maintaining a second, model-specific serving path next to a general one, or whether the headline tok/s justifies the extra deployment and maintenance burden. Sentiment was broadly positive but with real pushback: the top comment (u/TokenRingAI) argues AI hardware is too expensive to run at slow speeds – on his 2x Xeon Max, llama.cpp gets ~7 tok/s and SGLang ~12 on Qwen Flash Next, while a custom NUMA engine he built hits ~67 tg/s and ~900 pp/s – so the gap is large enough to matter. A skeptic (u/darktotheknight) frames it as an unsustainable fragmentation (the “Apache HTTP server” analogy ) that will consolidate. u/buttplugs4life4me makes the strongest case for the generalists: SGLang is already falling apart under model/hardware explosion (no GGUF support for the most common Qwen arch, no mixed non-block quant, MXFP8/MXFP4 unoptimized, no mixed KV cache). https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/
2. EXL3 / ExLlamaV3 on AMD Strix Halo (ROCm) for ~300B MoE models Several posts converge on ExLlamaV3 (EXL3) as the memory-efficient path for running 300B-class MoE models on 128 GB of unified-memory hardware. A new ROCm engine (Kyojin) targets Strix Halo specifically and fits two 300B MoE models on one 128 GB machine with low KLD. For your NVIDIA-only stack this is less directly actionable, but the EXL3 layer-mix quantization technique (mixing ~2.05 and 3.05 bpw tensors to balance context vs. KLD) is the technique to watch, and EXL3 weights are portable where the backend is supported. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/
3. Qwen3.8 Flash-Next and Qwen3.8-27B on constrained hardware (16 GB class) Two separate posts pushed Qwen3.8 Flash-Next 176B and Qwen3.8-27B onto 16 GB-class GPUs + SSD, using tiered scheduling (VRAM/RAM/SSD) and custom quants. Directly relevant if you are probing whether a 3090 (24 GB) or a laptop can serve these models without a multi-GPU rig. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/ https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/
(splitting because of post character limit, see below for the rest)
New models
Aleph-Alpha/Kolibri-1 – 78.1B total / 3.46B active, up to 1M context, Apache 2.0 A new open-weight English-German Mixture-of-Experts reasoning model from the German lab Aleph Alpha, open-sourced Oct 3, 2026 (Germany’s Reunification Day). 78.1B total params, ~3.46B active per token, native context up to 1,048,576 tokens, weights on Hugging Face under Apache 2.0. Trained on ~20T tokens (mostly English ~62.5%, German ~23.9%, code ~13.6%). Notable as one of the largest sovereign open-weight releases from Europe. Why it matters to you: the 3.46B-active MoE makes it a strong candidate to replace GPT-OSS at roughly two-thirds the size for German/English workloads, and the 1M context is a big draw. Commenters note the model is “basiclly worse than Qwen3.6-35B at twice the size” (u/Training_Visual6159) and that it is “comparable to Qwen 3.5 35B” (u/SnooPaintings8639), so treat it as a strong first attempt rather than a clear winner. At ~78B it likely needs a 96 GB VRAM build to run comfortably. As of this digest there were no community quant releases (no GGUF/INT4/NVFP4 quants posted), and a commenter asked whether llama.cpp supports the architecture yet – so do not plan a deployment until a serving path and quants exist. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/ Tech report: https://aleph-alpha.com/downloads/tech-report.pdf Model: https://huggingface.co/aleph-alpha/Kolibri-1
Qwen3.8 Flash-Next (~176B total, 6B active) – the recurring constrained-hardware test subject Not new today, but it is the model two of the five featured posts are benchmarking on 16 GB-class hardware. The MoE with 512 experts and a large n-gram/PLE table makes it unusually suited to tiered SSD offloading. See the “Most interesting posts” section for the two 16 GB data points.
bilibili “Index-Translate” – multilingual translation family on Qwen3.5 (lower signal) A small (15?) post announcing a multilingual translation model family based on Qwen3.5. Translation-focused, so limited relevance to your coding/agentic backend; flagging it only because it is a new Qwen3.5-based release. No quants or serving configs posted. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wxa1wr/bilibili_released_indextranslatea_a_multilingual/
Most interesting posts
1. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/ u/Yaniss916 reports two ~300B MoE models, each on ONE 128 GB AMD Strix Halo (Ryzen AI Max+ 395, gfx1151) running a new ExLlamaV3/ROCm engine called Kyojin. Concrete numbers from the post: GLM-5.3-Flash (99.7 GB) hits ~580 tok/s prefill (at 3.5K), 546 at 64K, 26-30 tok/s decode (MTP); MiMo-V2.6-Flash-MOPD (105 GB) hits ~650 tok/s prefill at 4K and up to 44 tok/s decode (speculative, code). Quality: KLD vs official FP8 of 0.151 / 0.0713 and top-1 agreement 89.3% / 92.0% respectively. The GLM pack mixes turboderp’s public 2.05 and 3.05 bpw EXL3 tensors with a custom layer mix. Includes a quickstart command (clone, ./build.sh, hf download, python tools/glm/serve.py --model ./glm-pack -c 131072 --num-draft 2, OpenAI-style API). Reproducible on Strix Halo hardware only – the author explicitly lists as “not measured yet” any GPU other than gfx1151, and the conversion pipeline stays private. Highly relevant as a technique reference (EXL3 layer-mix quant + ROCm) even though the target ha rdware is AMD, not your NVIDIA stack. Resources: Engine https://github.com/Yamz-Labs/kyojin . Weights https://huggingface.co/yamz-labs . Base fork https://github.com/vcruz305/exllamav3-amd
2. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/ u/fuzhongkai runs Qwen3.8 Flash-Next 176B on an RTX 3080 Laptop (16 GB VRAM) + 32 GB RAM + SSD using his open-source engine TensorSharp, which does MoE-aware scheduling across cache, VRAM, system RAM, and SSD. Reported results: ~11.09 tok/s decode vs Strata’s 10.24, but end-to-end 16.54 s vs Strata’s 62.15 s on his test (peak GPU 14.8 GB, OS working set 19.74 GiB). Reproducibility caveat (flagged in comments): the post does not name the quantization, and top comments ask for exactly that (u/DigitalguyCH: “very vague without specifying the quantization”; u/wizard_of_menlo_park: “Quantization?”). u/leonbollerup also says he gets ~80 tok/s decode and ~2200 tok/s prefill on the same card with Strata, which contradicts the post’s headline – so treat the end-to-end speedup as single-source and unverified. This is directly relevant to whether a 24 GB 3090 can serve the 176B Flash-Next via SSD tiering, but verify the quant and benchmark method before trusting it. Resources: https://github.com/zhongkaifu/TensorSharp . Model docs https://github.com/zhongkaifu/TensorSharp/blob/main/docs/models/qwen38-flash-next.md
3. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/ u/roofkid built NInfer 4080 to run ISTA-DASLab Qwen3.8-27B-GSQ at 100k context on an RTX 4080 16 GB, claiming up to 2720 tok/s prefill and 262 tok/s generation. Part of the NInfer 5090/4090/3090 family of from-scratch engines. Reproducibility caveat: the post is a project announcement; the headline numbers (2720 pp / 262 gen) are the author’s max-observed, and a top comment (u/Pyrolistical) notes these custom engines often compromise with a quantized KV cache and that his own roofline analysis of llama.cpp on an R9700 found <5% decode improvement left on the table – so the gap between custom and general engines is hardware/kernel-specific, not universal. Comments are mostly “how do I get this on a 3080/5080/4070 Ti” requests, so portability to other cards is unconfirmed. Relevant to you primarily as another data point in the overfit-engine trend and for the ISTA-DASLab 3-bit GSQ quant on Qwen3.8-27B. Resources: https://github.com/roofkid/ninfer-4080 . Quant https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ . ByteShape Qwen3.8-27B reference https://byteshape.com/blogs/Qwen3.8-27B/
4. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/ u/carteakey’s discussion post on the overfit-engine trend (see topics above). 315?, 199 comments. No single reproducible config, but the most substantive community engineering observations of the day: SGLang’s current gaps (u/buttplugs4life4me), the Xeon-Max NUMA speed gap (u/TokenRingAI), and a practitioner who “vibe-coded” his own engine because upstream PRs were ignored (u/wishstudio). Worth reading for the state of general vs. narrow runtimes rather than as a runnable config.
5. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/ The Kolibri-1 announcement (see New models above). 510?, 153 comments. Not a reproducible-serving post (no quants yet), but the most-discussed new model of the day and directly relevant to a European/sovereign model pipeline and as a GPT-OSS-sized MoE candidate.
Useful resources discovered
https://github.com/Yamz-Labs/kyojin Found in the Kyojin post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/). An ExLlamaV3-based ROCm engine with a quickstart that serves GLM-5.3-Flash / MiMo with an OpenAI-style API. Matters to you as the reference implementation for EXL3 layer-mix quantization and for the MoE-expert offloading pattern, though it is ROCm/Strix-Halo specific and the conversion pipeline is private. Actionable if you have AMD hardware or want to study the quant technique; not directly runnable on your NVIDIA stack as-is.
https://huggingface.co/turboderp/Qwen3.8-Flash-Next-exl3 Found in a Kyojin comment (u/MarkoMarjamaa, https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/?context=3#pdmp0wf). Public EXL3 weights for Qwen3.8 Flash-Next. The comment notes “with exl3 it would be possible to run Q5-quality quant in Q4 memory.” Directly relevant to your Qwen3.8-27B/Flash-Next deployments if you adopt ExLlamaV3; immediately actionable for EXL3 users.
https://github.com/zhongkaifu/TensorSharp Found in the 176B-on-3080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/). An open-source inference engine doing MoE-aware tiered scheduling (VRAM/RAM/SSD) so a model far larger than VRAM and RAM stays usable. Relevant if you want to serve a 176B MoE on a 24 GB 3090 via SSD; the model-specific doc page (…/docs/models/qwen38-flash-next.md) has the exact launch path. Actionable, but the author has not published the quant used in his benchmark – verify before relying on the tok/s.
https://github.com/roofkid/ninfer-4080 Found in the Ninfer-4080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/). A from-scratch engine running ISTA-DASLab Qwen3.8-27B-GSQ at 100k context on a 16 GB 4080, with a runnable script (…/scripts/run-ninfer-4080.bat). Actionable for 16 GB-class Ampere cards; portability to other cards is unconfirmed and the KV-cache quant is a known tradeoff.
https://huggingface.co/aleph-alpha/Kolibri-1 (+ tech report https://aleph-alpha.com/downloads/tech-report.pdf) Found in the Kolibri-1 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/). Official 78.1B / 3.46B-active English-German MoE weights under Apache 2.0 with a 1M-context spec and a full tech report. Actionable for evaluation once a serving path/quants exist; not deployable on your stack today without a backend and quant.
https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ Found in the Ninfer-4080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/). A 3-bit GSQ quant of Qwen3.8-27B that the engine uses. Relevant if you are tracking sub-4-bit quants for 27B-class models; 3-bit quality for agentic work is unverified here – treat as an extreme-quant data point, not a recommended serving quant.
Interesting anecdotes
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdpighm u/Toxaris71 reports ~11 tok/s decode / 30 tok/s prefill on Qwen 3.8 Flash IQ3_XXS (8 GB 3070 Ti + 32 GB DDR4), and 32 tok/s decode / 80 tok/s prefill with IQ2_0, saying it is “better quality than Orinth 1.5 35B 4-bit” and much faster than Qwen 3.8 27B IQ4_XS (~2.8 tok/s decode). Interesting as a low-VRAM 3070 Ti comparison point for the Flash-Next/27B quants, but no launch command, backend version, or benchmark harness is given – treat as anecdotal.
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdofln4 u/leonbollerup claims ~80 tok/s decode and ~2200 tok/s prefill on the same 16 GB 3080 using Strata, which contradicts the post author’s ~11 tok/s decode. Useful as an upper-bound comparison for Flash-Next on 16 GB, but no quant or config is provided – anecdotal and in direct tension with the post.
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdq00tm u/Independent_Grade612 reports ~450 prefill / ~30 decode on a 12 GB RTX 3500 Ada + 64 GB RAM using Strata with Swift 1.5 Q2XS. Another single-number 12 GB data point; no reproducible config – anecdotal.
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/?context=3#pdocpvd u/TokenRingAI reports ~7 tok/s on llama.cpp and ~12 on SGLang for Qwen Flash Next on his 2x Xeon Max, versus ~67 tg/s and ~900 pp/s on a custom NUMA engine. Notable as the most concrete performance-gap claim in the overfit-engine thread, supporting the case that narrow engines can dramatically outperform general ones on specific hardware. Still anecdotal (no command/config), but the numbers are specific enough to be worth reproducing.
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/?context=3#pdlfi5y u/FullstackSensei estimates Kolibri-1’s training cost at ~$4M (20T tokens on a 768x B300 cluster over 4 weeks at ~$8/GPU-hr). Directionally interesting for understanding training-cost trends, but it is a single commenter’s estimate, not a published figure – treat as a rough back-of-envelope, not a fact.
Well, that’s super wordy for my taste, but that’s the beauty of local, you can change things to your taste. I’ll probably go Gemma 4 12B or 31B (Qwen 3.8 27B is great for code but I like Gemma for text munging), and am AMD based, but hermes is probably a good tool.
Did you just mod the Reddit Reading skill to use redlib? Why not just use the skill raw (presumably with the OAuth)?
Thanks again.
this channel https://youtu.be/lHmZoRHMZyM
Nah, I find text to have a much higher bandwidth (ironically), waiting for them to get to the point or finish something I already know drives me nuts. Videos are for something that needs the medium, instructions for disassembling something for example.
Hacker News will be a good aggregator of the top AI stories and then you can see what blogs pop up regularly to add directly to your feed
you can subscribe to !hackernews@lemmy.bestiver.se and then cross-post to here when appropriate






