I currently run the top 8 or so per day from r/LocalLLama through my RSS reader. Actually going to reddit is eww, so I want it to be worthwhile. I usually get enough to keep roughly up to date just with title and text / graphs and only go if it looks really interesting.

With the upcoming demise of reddit RSS, I’m wondering if people have recommended (hopefully RSS friendly) places roughly equivalent. I guess I could filter HN which may be useful for many tech interests, but I don’t know what I don’t know, so I thought I’d pose the question to the crowd.

  • MalReynolds@slrpnk.netOP
    link
    fedilink
    English
    arrow-up
    3
    ·
    2 days ago

    Now there’s an interesting idea. Ooh, redlib even has a quadlet definition out of the box, don’t see that every day. I may well give that a spin, probably tack a rss feed on the end as well (creature of habit). Thanks.

    • Dran@lemmy.world
      link
      fedilink
      English
      arrow-up
      5
      ·
      2 days ago

      if you actually intend to replicate my setup: [qwen3.8:27:nvfp4 --> vllm (concurrency efficiency)] --> [open-webui (api proxy)] --> [hermes-agent --> hermes-webui]

      […] denotes container boundaries.

      The redlib instance also runs in a container on the same host. A skill teaches hermes-agent how to use redlib to open a subreddit and a separate instruction markdown file teaches it how to build the digest I want. a “cron” (hermes cron, not system cron) runs the job once a day, and the final step of the digest instructions is to email me the digest.

      This was today’s, for example:

      LocalLLaMA Daily Digest – 2026-10-04

      Curated for a [REDACTED] LLM-backend sysadmin running vLLM + Open-WebUI on RTX 3090s and RTX PRO 6000 Blackwell. Focus: quantization (INT4/NVFP4/GGUF/EXL3/EXL2), inference serving, VRAM-limited local deployment, and reproducible configs.


      Most frequently discussed topics

      1. The rise of narrow “overfit” inference engines This was the loudest infrastructure thread today. A wave of deliberately non-general runtimes – Strata, Ninfer, DwarfStar, Splash, llamAmpere, gufo, Kyojin, TensorSharp – give up llama.cpp/vLLM’s generality to squeeze maximum throughput out of a handful of models on a single hardware family (often one GPU or the Strix Halo APU). For a vLLM/Open-WebUI operator, the question is whether it is worth maintaining a second, model-specific serving path next to a general one, or whether the headline tok/s justifies the extra deployment and maintenance burden. Sentiment was broadly positive but with real pushback: the top comment (u/TokenRingAI) argues AI hardware is too expensive to run at slow speeds – on his 2x Xeon Max, llama.cpp gets ~7 tok/s and SGLang ~12 on Qwen Flash Next, while a custom NUMA engine he built hits ~67 tg/s and ~900 pp/s – so the gap is large enough to matter. A skeptic (u/darktotheknight) frames it as an unsustainable fragmentation (the “Apache HTTP server” analogy ) that will consolidate. u/buttplugs4life4me makes the strongest case for the generalists: SGLang is already falling apart under model/hardware explosion (no GGUF support for the most common Qwen arch, no mixed non-block quant, MXFP8/MXFP4 unoptimized, no mixed KV cache). https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/

      2. EXL3 / ExLlamaV3 on AMD Strix Halo (ROCm) for ~300B MoE models Several posts converge on ExLlamaV3 (EXL3) as the memory-efficient path for running 300B-class MoE models on 128 GB of unified-memory hardware. A new ROCm engine (Kyojin) targets Strix Halo specifically and fits two 300B MoE models on one 128 GB machine with low KLD. For your NVIDIA-only stack this is less directly actionable, but the EXL3 layer-mix quantization technique (mixing ~2.05 and 3.05 bpw tensors to balance context vs. KLD) is the technique to watch, and EXL3 weights are portable where the backend is supported. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/

      3. Qwen3.8 Flash-Next and Qwen3.8-27B on constrained hardware (16 GB class) Two separate posts pushed Qwen3.8 Flash-Next 176B and Qwen3.8-27B onto 16 GB-class GPUs + SSD, using tiered scheduling (VRAM/RAM/SSD) and custom quants. Directly relevant if you are probing whether a 3090 (24 GB) or a laptop can serve these models without a multi-GPU rig. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/ https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/

      (splitting because of post character limit, see below for the rest)