I currently run the top 8 or so per day from r/LocalLLama through my RSS reader. Actually going to reddit is eww, so I want it to be worthwhile. I usually get enough to keep roughly up to date just with title and text / graphs and only go if it looks really interesting.

With the upcoming demise of reddit RSS, I’m wondering if people have recommended (hopefully RSS friendly) places roughly equivalent. I guess I could filter HN which may be useful for many tech interests, but I don’t know what I don’t know, so I thought I’d pose the question to the crowd.

  • terabyterex@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    ·
    10 hours ago

    from what i am hearing and i really really really dont want to create an scvount but X (teitter). ftom what i gear, its a good resource. people are sharing good things

  • keepthepace@tarte.nuage-libre.fr
    link
    fedilink
    English
    arrow-up
    8
    ·
    17 hours ago

    Reddit is eventually going to die of enshitification. The best way to survive it is to make this alternative work and promote it regularly (in a non spammy way) on reddit.

    • ThorrJo@lemmy.sdf.org
      link
      fedilink
      English
      arrow-up
      3
      ·
      9 hours ago

      also not OP but I am interested in local models usable on low-end hardware, models for tasks like document summarizing and text rewording, fuzzy search of various types of local data, and search agents such that I can ask it a vague natural-language question and it runs 17 web searches and comes back with a summary of what it found.

    • EnsignWashout@startrek.website
      link
      fedilink
      English
      arrow-up
      4
      ·
      17 hours ago

      I’m not the OP, but I’m trying to figure out if local models are always painfully slow or if I’m missing something obvious in my tuning.

      Stable Diffusion can whip up a picture in less time on the same hardware, than lama.cpp takes to decide to call an MCP function.

      It seems like I must be missing something in my lama.cpp setup, but none of the guides I’ve read have clued me in to what I’ve done wrong.

      Ollama performs similarly poorly on the same harsware, so I’ve probably managed to make the se mistake(s) at least twice.

      Anyway, that’s the main thing I’m reading along for. Trying to increase my understanding until I catch my own mistakes.

  • Dran@lemmy.world
    link
    fedilink
    English
    arrow-up
    7
    ·
    21 hours ago

    I do something similar to you, except I use redlib as a backend and a local qwen3.8:27b as an aggregator to build me a daily digest every day.

    If we innovate a little we can survive the death of rss

    • PetteriPano@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      16 hours ago

      I have my hermes agent run xvnc with chrome & playwright to navigate old.reddit.com with an old throwaway account. It has created a little script to reduce the DOM to an LLM-friendly size.

      It also digests the consensus of the comments in the top posts.

    • MalReynolds@slrpnk.netOP
      link
      fedilink
      English
      arrow-up
      3
      ·
      20 hours ago

      Now there’s an interesting idea. Ooh, redlib even has a quadlet definition out of the box, don’t see that every day. I may well give that a spin, probably tack a rss feed on the end as well (creature of habit). Thanks.

      • Dran@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        17 hours ago

        if you actually intend to replicate my setup: [qwen3.8:27:nvfp4 --> vllm (concurrency efficiency)] --> [open-webui (api proxy)] --> [hermes-agent --> hermes-webui]

        […] denotes container boundaries.

        The redlib instance also runs in a container on the same host. A skill teaches hermes-agent how to use redlib to open a subreddit and a separate instruction markdown file teaches it how to build the digest I want. a “cron” (hermes cron, not system cron) runs the job once a day, and the final step of the digest instructions is to email me the digest.

        This was today’s, for example:

        LocalLLaMA Daily Digest – 2026-10-04

        Curated for a [REDACTED] LLM-backend sysadmin running vLLM + Open-WebUI on RTX 3090s and RTX PRO 6000 Blackwell. Focus: quantization (INT4/NVFP4/GGUF/EXL3/EXL2), inference serving, VRAM-limited local deployment, and reproducible configs.


        Most frequently discussed topics

        1. The rise of narrow “overfit” inference engines This was the loudest infrastructure thread today. A wave of deliberately non-general runtimes – Strata, Ninfer, DwarfStar, Splash, llamAmpere, gufo, Kyojin, TensorSharp – give up llama.cpp/vLLM’s generality to squeeze maximum throughput out of a handful of models on a single hardware family (often one GPU or the Strix Halo APU). For a vLLM/Open-WebUI operator, the question is whether it is worth maintaining a second, model-specific serving path next to a general one, or whether the headline tok/s justifies the extra deployment and maintenance burden. Sentiment was broadly positive but with real pushback: the top comment (u/TokenRingAI) argues AI hardware is too expensive to run at slow speeds – on his 2x Xeon Max, llama.cpp gets ~7 tok/s and SGLang ~12 on Qwen Flash Next, while a custom NUMA engine he built hits ~67 tg/s and ~900 pp/s – so the gap is large enough to matter. A skeptic (u/darktotheknight) frames it as an unsustainable fragmentation (the “Apache HTTP server” analogy ) that will consolidate. u/buttplugs4life4me makes the strongest case for the generalists: SGLang is already falling apart under model/hardware explosion (no GGUF support for the most common Qwen arch, no mixed non-block quant, MXFP8/MXFP4 unoptimized, no mixed KV cache). https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/

        2. EXL3 / ExLlamaV3 on AMD Strix Halo (ROCm) for ~300B MoE models Several posts converge on ExLlamaV3 (EXL3) as the memory-efficient path for running 300B-class MoE models on 128 GB of unified-memory hardware. A new ROCm engine (Kyojin) targets Strix Halo specifically and fits two 300B MoE models on one 128 GB machine with low KLD. For your NVIDIA-only stack this is less directly actionable, but the EXL3 layer-mix quantization technique (mixing ~2.05 and 3.05 bpw tensors to balance context vs. KLD) is the technique to watch, and EXL3 weights are portable where the backend is supported. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/

        3. Qwen3.8 Flash-Next and Qwen3.8-27B on constrained hardware (16 GB class) Two separate posts pushed Qwen3.8 Flash-Next 176B and Qwen3.8-27B onto 16 GB-class GPUs + SSD, using tiered scheduling (VRAM/RAM/SSD) and custom quants. Directly relevant if you are probing whether a 3090 (24 GB) or a laptop can serve these models without a multi-GPU rig. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/ https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/

        (splitting because of post character limit, see below for the rest)

    • MalReynolds@slrpnk.netOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      6 hours ago

      Nah, I find text to have a much higher bandwidth (ironically), waiting for them to get to the point or finish something I already know drives me nuts. Videos are for something that needs the medium, instructions for disassembling something for example.

  • blueduck@piefed.social
    link
    fedilink
    English
    arrow-up
    3
    ·
    21 hours ago

    Hacker News will be a good aggregator of the top AI stories and then you can see what blogs pop up regularly to add directly to your feed