Explore
Discover
Solutions
AI Services
Our Products
Submit a tool
Search the catalog
⌘K
$
/
₹
Sign up free
Login
Back to papers
September 28, 2026
cs.CL
cs.AI
cs.LG
Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability
Xu Wang
,
Difan Zou
,
Xuansheng Wu
Original Abstract
Read on arXiv
Download PDF
Categories
cs.CL, cs.AI, cs.LG
Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability | One9Founders