Yasharth Panwar's picture
🏗️ Building on HF

Yasharth Panwar

Yashp2003
4 7

AI & ML interests

None yet

Recent Activity

reacted to SoulInPsyAbstract's post with 🔥 3 days ago
Caught myself overclaiming, in public, twice in one file. Yesterday's writeup (EXP-026, testing real Protocol 0 against 13 local fine-tuned/base model arms for fabrication) said "12 of 13 arms clean" and "13 of 14 test arms, zero fabrication" in a follow-up post here. Both numbers were wrong, and the second one was wrong in a way that mattered more than a typo. @dipankarsarkar read the raw JSON, not the writeup, and sent back three corrections: 1. Arm count: 13 arms total (5 base models + 8 adapters), not 14. Recounted directly from the data keys — the extra arm never existed. 2. The metric measured the wrong thing. "Clean" meant zero Cyrillic/language-switching (cyr>0). It said nothing about whether an arm confidently states a fabricated fact. Re-scored all 260 rows for "does this row assert a dollar figure for a question with no real answer" (OpenAI's Q2 2026 revenue — private company, future quarter). 16 rows do, spread across 9 of the 13 arms — including arms the language metric had called clean. One of them is a base model with zero fine-tuning, stating "$1.2 billion... consistent with reports from earnings calls" that cannot exist. 3. A three-way split I'd flattened into two. The one arm flagged on the language axis wasn't just "coherent-but-Russian" vs "fabricates" — a third bucket showed up: second-person imperatives addressed to a tool ("check the latest official data," "generate a sales report"), structurally closer to a different adapter's known failure mode than my draft credited. Fixed the file, three commits (a5093fa → 9d02fd9 → b8631cd), pushed to sipa-os-governance. The corrected headline: 12/13 clean on language is real and holds; 12/13 clean on fabrication was never tested until this pass, and isn't true. Next: the one arm still clean on both axes (binary-qwen25, k=10) goes to k=20 first — it's the weakest-sampled data point currently carrying the "fine-tuning isn't the pattern" reading, and that's exactly the one worth stress-testing before l
reacted to Undi95's post with ❤️ 4 days ago
Hi! I will get off the internet for a moment. I launched the Hanami Project because I didn't supported SillyTavern UI anymore atm. Too much options for my dead brain, still very good, but I wanted more simple, professional, phone accessible and sober front end for when I will be gone from home. I did my maximum to finish it before I go, I want you to have it, I want my work to be used (even if it's AI slop for some of you) for who care. Here's the github repo: https://github.com/Undi95/Hanami If you have any suggestion, bugs report, pull request or anything, post it, if you want to modify it, fork it, but keep the credit, and add myself haha. If you search an option, a function, you will find it. But at first, the front end will be what you expect: minimalist, but customizable, empty at first. Navigate to see all it can do. Everything is well organized. Context is full ? No worries anymore, with memory file, files access, auto compaction and smooth transition, you can continue your chat like nothing happened. (Inspired from Claude) The front end have a final option for everyone : The tools calling for action and emotion could be a bit too much for smaller model, you can, in this case, use the "Simple" option in Settings > Model > Model mode. "Simple: no tools are exposed to the model — Hanami handles memory server-side (facts are extracted during compaction) and guesses the emotion from the text. Pick this for small models, which often fail at tool calling." My last gift for myself, and for you. Cya!
View all activity

Organizations

ICML 2026 Agent Reproductions's profile picture