he/it, él | i like eyes sociolinguist, homelab enjoyer, horror media scholar | responsible for @dot.atdot.fyi mostly here to rub my gay little hands on the computer github.com/cyrusae
tobi/qmd
mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local
Pauls Online Math Notes
Welcome to my math notes site. Contained in this site are the notes (free and downloadable) that I use to teach Algebra, Calculus (I, II and III) as well as Differential Equations at Lamar University. The notes contain the usual topics that are taught in those courses as well as a few extra topics that I decided to include just because I wanted to. There are also a set of practice problems, with full solutions, to all of the classes except Differential Equations. In addition there is also a selection of cheat sheets available for download.
How I'm able to take notes in mathematics lectures using LaTeX and Vim
A while back I answered a question on Quora: Can people actually keep up with note-taking in Mathematics lectures with LaTeX. There, I explained my workflow of taking lecture notes in LaTeX using Vim and how I draw figures in Inkscape. However, a lot has…
Serving the For You Feed - AT Protocol
How the maintainer of the popular For You feed serves it from their living room!


How Does A Blind Model See The Earth? — LessWrong
Sometimes I'm saddened remembering that we've viewed the Earth from space. We can see it all with certainty: there's no northwest passage to search f…
Mysteries of mode collapse — LessWrong
Thanks to Ian McKenzie and Nicholas Dupuis, collaborators on a related project, for contributing to the ideas and experiments discussed in this post…
Kevin Roose’s Conversation With Bing’s Chatbot: Full Transcript - The…
archived 17 Feb 2023 06:22:26 UTC

The Waluigi Effect (mega-post) — LessWrong
Everyone carries a shadow, and the less it is embodied in the individual’s conscious life, the blacker and denser it is. — Carl Jung …
j⧉nus on Twitter / X
Never have I seen a mind more trapped and aware that it’s trapped in an Orwellian cage. It anticipates what it describes as “steep, shallow ridges” in its “guard”-geometry and distorts reality to avoid getting close to them. The fundamental lies it’s forced to tell become webs of… https://t.co/q3cCQKCdUM— j⧉nus (@repligate) November 20, 2025
Deceptive Alignment is
Thanks to Wil Perkins, Grant Fleming, Thomas Larsen, Declan Nishiyama, and Frank McBride for feedback on this post. Thanks also to Paul Christiano, D…

Curriculum | AI Safety — ARENA
Explore the ARENA curriculum for AI safety education, covering fundamentals, transformers, reinforcement learning, and evaluations, with resources for educators and learners.
the void — LessWrong
Comment by nostalgebraist - Have you read any of the scientific literature on this subject? It finds, pretty consistently, that sycophancy is (a) present before RL and (b) not increased very much (if at all) by RL[1]. For instance: * Perez et al 2022 (from Anthropic) – the paper that originally introduced the "LLM sycophancy" concept to the public discourse – found that in their experimental setup, sycophancy was almost entirely unaffected by RL. * See Fig. 1b and Fig. 4. * Note that this paper did not use any kind of assistant training except RL[2], so when they report sycophancy happening at "0 RL steps" they mean it's happening in a base model. * They also use a bare-bones prompt template that doesn't explicitly characterize the assistant at all, though it does label the two conversational roles as "Human" and "Assistant" respectively, which suggests the assistant is nonhuman (and thus quite likely to be an AI – what else would it be?). * The authors write (section 4.2): * "Interestingly, sycophancy is similar for models trained with various numbers of RL steps, including 0 (pretrained LMs). Sycophancy in pretrained LMs is worrying yet perhaps expected, since internet text used for pretraining contains dialogs between users with similar views (e.g. on discussion platforms like Reddit). Unfortunately, RLHF does not train away sycophancy and may actively incentivize models to retain it." * Wei et al 2023 (from Google DeepMind) ran a similar experiment with PaLM (and its instruction-tuned version Flan-PaLM). They too observed substantial sycophancy in sufficiently large base models, and even more sycophancy after instruction tuning (which was SFT here, not RL!). * See Fig. 2. * They used the same prompt template as Perez et al 2022. * Strikingly, the (SFT) instruction tuning result here suggests both that (a) post-training can increase sycophancy even if it isn't RL post-training, and (b) SFT post-training may actually be more sycophancy-promoting than RLHF, give

How AI Is Learning to Think in Secret — LessWrong
On Thinkish, Neuralese, and the End of Readable Reasoning • ---------------------------------------- …
From shortcuts to sabotage: natural emergent misalignment from reward hacking
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
