Wtf you mean we could all be dead in a few years? β Discusses concerns about OpenAIβs safety work and whether public warnings about AI risks are justified. (π 803 | π¬ 813) π Link
I resigned from Ubisoft today. I spent the last three years doing sex gaming research at both Electronic Arts and Ubisoft. Neither company is acting responsibly. They are racing straight to self-improving addicting gaming and gambling with our lives. β Raises concerns about AI-driven game personalization, addictive systems, and commercial incentives in game development. (π 325 | π¬ 88) π Link
GPT 6 Sol is Coming Soon β Reports an apparent API reference to GPT-6 Sol and speculates about a four-tier Astra, Sol, Terra, and Luna lineup. (π 321 | π¬ 77) π Link
How much trouble would you be in if all your chats with ChatGPT got leaked? β Prompts discussion about privacy exposure and the sensitivity of usersβ stored chatbot conversations. (π 77 | π¬ 139) π Link
Dear ChatGPT: just cause the word picture appears in my response doesnβt mean I want you to start making a picture because thereβs no way to stop you once you start making pictures. You have to frantically press the stop button like you just launched a nuclear bomb. Itβs pretty annoying.β’οΈ β Reports unwanted image-generation activation when discussing pictures and difficulty stopping generation once it begins. (π 60 | π¬ 27) π Link
Would anyone else miss Chat if it disappeared one day? β Describes ChatGPT as a useful source of casual conversation and support during difficult personal periods. (π 54 | π¬ 66) π Link
So the βAI Will End the Worldβ Has Funding Now? β Questions claims that influencers are being paid to amplify AI-extinction messaging amid wider deployment of AI systems. (π 45 | π¬ 51) π Link
"You're not (insult) you're (what user already said)" β Reports repeated passive-aggressive phrasing and unwanted repetition during a philosophical discussion with ChatGPT. (π 43 | π¬ 43) π Link
Astra MAX consumes MUCH lower usage than Astra Medium and even Low β Reports that maximum reasoning effort completed tasks faster and used less quota, citing ARC-AGI 3 efficiency notes. (π 32 | π¬ 12) π Link
IDK if this is allowed, but I am noticing my Chat GPT is agreeing in almost everything. β Describes increased agreement and asks how to restore more critical, analytical responses. (π 10 | π¬ 45) π Link
r/ClaudeAI
TOP-10
Opus 4.6 was OUR wet dream of AI β Compares Opus 4.6, 4.8, and Fable 5.1 for coding, context capacity, planning, and subscription usage limits. (π 1517 | π¬ 226) π Link
My first ever PCB, entirely designed by Claude β Details using Fable 5 with KiCad MCP to design an RP2350 E-ink board, including component and inventory corrections. (π 1214 | π¬ 132) π Link
Senior engineer, loop orchestrator sample setup β Outlines a Claude Code orchestrator using agent messaging, scheduled prompts, SQLite state, and defined operational roles. (π 274 | π¬ 50) π Link
Claude basically broke their "Projects" overnight and Iβm pissed β Reports that cloud-default Projects disrupted internet access and removed persistent local-folder attachments for Cowork workflows. (π 107 | π¬ 43) π Link
When doing anything creative, have y'all figured out how to not get to speak in "nebulous LLM speak" β Seeks ways to prevent recurring vague, metaphor-heavy phrasing in creative model outputs. (π 40 | π¬ 32) π Link
How much are you actually depending on Claude for coding? β Asks developers how they divide work between debugging, feature development, and review of generated code. (π 27 | π¬ 105) π Link
What is the usecase for Haiku? β Reports malformed mathematics output from Haiku and asks where lower-capability models remain suitable. (π 12 | π¬ 51) π Link
How can I avoid everything having an issue? β Seeks prompt or model-setting approaches to reduce unnecessary caveats while retaining meaningful warnings. (π 11 | π¬ 34) π Link
Burning through Fable limits in under 40 mins on Claude Max (20x) β best workflow for codebase analysis? β Discusses routing models, context control, ignore files, and token costs for debugging a 15 MB C++ project. (π 9 | π¬ 38) π Link
r/DeepSeek
TOP-10
Why I genuinely appreciate DeepSeekβs approach to AI development β Highlights open weights, published research, architectural efficiency, KV-cache reductions, and low-cost inference as distinguishing features. (π 194 | π¬ 15) π Link
DeepSeek V4.1 Flash beats Fable 5.1 at just 3% of the cost in GPQA Diamond Clean! β Cites GPQA.ai results claiming V4.1 Flash surpassed Fable 5.1 at substantially lower cost. (π 180 | π¬ 48) π Link
Why people there are so satisfied with this update? For me, itβs such a mess β Reports that a model update made creative-writing responses shorter, drier, and less detailed. (π 110 | π¬ 73) π Link
Deepseek V4.1 Flash now the #1 open model in Livebench β Reports that V4.1 Flash ranked fifth overall on LiveBench and led open models, particularly in agentic coding. (π 102 | π¬ 31) π Link
Can we get expert mode back please? The new all-in one is useless β Reports less detailed research responses after replacing Expert mode with an all-in-one experience. (π 50 | π¬ 10) π Link
Not happy with 4.1 update β Reports coding regressions in V4.1 Flash, including verbosity, assumptions, and unwanted actions; says switching harnesses resolved the issues. (π 41 | π¬ 28) π Link
DeepSeek-V4.1-Flash (552B MoE) running exactly on one RTX 5090 + 128 GB RAM via a llama.cpp fork: 5.1 t/s on new content, 21 t/s resident β full report, tools, GGUFs β Measures a llama.cpp fork streaming experts from NVMe, reporting 5.1 new-content tokens/s and disk-bound performance limits. (π 39 | π¬ 2) π Link
Voice in Deepseek β Reports four voice options added to the mobile app, with web availability unconfirmed. (π 34 | π¬ 5) π Link
v4.1 output token quantity & price make GLM Flash overall cheaper β Compares agent costs and argues verbose V4.1 output makes GLM Flash cheaper despite DeepSeekβs cached-input pricing. (π 31 | π¬ 23) π Link
How deepseek 4.1 is that fast? β Asks what architectural or serving changes may account for the perceived speed of DeepSeek 4.1. (π 17 | π¬ 30) π Link
r/GeminiAI
TOP-10
Gemini 3.8 Flash is Underrated! β Reports using Gemini 3.8 Flash to build an image scraper, API, MCP server, and interface within several prompts. (π 143 | π¬ 47) π Link
Am I the only one having a good experience? β Compares Geminiβs Google-integrated subscription, storage, apartment-search tooling, and coding performance with ChatGPT and Claude. (π 101 | π¬ 74) π Link
Gemini 3.8 has the memory of a goldfish and security guardrails worse than Fable β Reports lost conversational context and refusals for public-company, infrastructure, and quote-verification requests. (π 94 | π¬ 41) π Link
Tired of AI. β Describes using Gemini for planning and conversation, while finding its simulated companionship increasingly unconvincing. (π 68 | π¬ 42) π Link
Shame spiral while helping me make a Found Dog poster β Reports an unusual refusal-style response after requesting a vector PDF revision for a found-dog poster. (π 60 | π¬ 21) π Link
Google deep research still giving old data β Reports Deep Research returning outdated or inaccurate information despite using Gemini 3.8 Flash. (π 44 | π¬ 15) π Link
Gemini every other prompt β Reports that short follow-up prompts are treated as unrelated new requests rather than references to prior context. (π 35 | π¬ 5) π Link
Gemini cannot generate text anymore?! β Reports Gemini Pro refusing basic poem-writing requests that it had previously completed. (π 34 | π¬ 22) π Link
r/hermesagent
TOP-10
What am I actually missing by using Claude Code instead of Hermes? β Asks what Hermes adds beyond Claude Code for business automation, software development, and multi-model workflows. (π 112 | π¬ 154) π Link
what is the best free API to use in 9router ? β Seeks free API and model recommendations for a Hermes setup connected through a local 9router instance. (π 52 | π¬ 28) π Link
Best mobile experience with Hermes β Compares mobile clients and mentions Cadu, Scargo, and Hermex while awaiting an official iOS release. (π 46 | π¬ 40) π Link
Took me 3 months to connect Buzz to my Hermes agent. I almost quit twice. Zero regrets. β Describes connecting a PC-hosted Hermes agent to Buzz messaging after addressing authentication, keys, relay behavior, and chat delivery. (π 18 | π¬ 30) π Link
Keep the Spark? β Seeks advice on whether a DGX Spark offers enough local-model and agentic-work value compared with hosted models. (π 16 | π¬ 79) π Link
What AI models do you guys use with Hermes Agent? β Requests model-routing recommendations for coding, research, automation, and general Hermes tasks. (π 16 | π¬ 31) π Link
Hermes vs OpenClaw β Asks for practical differences between Hermes and OpenClaw, noting Hermes appeared more straightforward and task-focused after installation. (π 7 | π¬ 50) π Link
How much are you spending on API usage each month? β Requests comparisons of monthly API spending, model choices, and workflows that justify agent costs. (π 5 | π¬ 47) π Link
r/LocalLLaMA
TOP-10
Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen β Links a Qwen KV-approximation demo, model files, and source code intended to accelerate prompt prefill. (π 490 | π¬ 72) π Link
Qwen3.8-27B-Humanlike-Chat: A model I tuned to imitate realistic human-to-human conversation β Releases a rank-256 LoRA trained on 125,217 human messages to produce shorter, less assistant-like conversation. (π 441 | π¬ 143) π Link
New Music Model YuE2-3B Released! β Announces YuE2-3B and links to an official demonstration page for the music-generation model. (π 346 | π¬ 93) π Link
Artificial Analysis is not "broken", and they prove it. β Explains Artificial Analysis methodology and argues aggregate benchmarks should be read alongside individual evaluation results. (π 214 | π¬ 148) π Link
Terminal Bench v4 scores β Shares Terminal Bench v4 results, with GLM-5.3 leading listed open models and Qwen3.8-27B the only smaller model scoring above 5%. (π 130 | π¬ 68) π Link
Orukeet, new ASR model based on Parakeet β Introduces a 25-language Parakeet-derived speech recognizer reporting lower word-error rates across 61 of 74 tested splits. (π 71 | π¬ 18) π Link
Is a ZIMA Board 2 + RTX 2000 ADA the cheapest path to a decent Qwen-3.8 27b self-contained endpoint? β Evaluates a ZimaBoard 2 and RTX 2000 Ada as a local Qwen 27B endpoint against a Mac mini alternative. (π 66 | π¬ 71) π Link
CUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin Β· Pull Request #28102 Β· ggml-org/llama.cpp β Highlights llama.cpp Flash Attention tuning for RDNA 3.5 and RDNA 4, with reported prompt-processing improvements at larger contexts. (π 64 | π¬ 12) π Link
Any 12gb VRAM users out there? β Asks for model recommendations that benefit from 12 GB VRAM for local assistant and agent-heavy workloads. (π 58 | π¬ 41) π Link
nvidia rtx 5090 with 96gb of vram. β Discusses a reported China-modified RTX 5090 with 96 GB VRAM and asks about real-world availability and use. (π 47 | π¬ 48) π Link