Essays and shorter notes, grouped by what they're about —
mostly AI safety. Filter by category below. Originally posted on LessWrong.
Latest
noteJun 20, 20262 karmaAdapting research plans on the flyResearch plans should be hypotheses, not contracts. Update them as evidence arrives while preserving a clear account of why the direction changed.
essayJun 19, 202637 karmaThe one-week sprintA one-week deadline can turn an ambitious, underspecified project into a concrete test of what matters. The constraint rewards decisive scoping, fast feedback, and finishing.
essayJun 18, 202681 karmaYour Model Organisms Might Be FriedTraining model organisms on synthetic documents can accidentally teach them that they are inside an experiment. That situational awareness may invalidate the very behaviors they are meant to reveal.
noteJun 7, 202681 karmaFake prerequisitesMany prerequisites are social stories rather than real dependencies. Test whether the supposedly necessary step actually blocks the work before spending months on it.
noteJun 6, 202615 karmaWhen to deduplicate work with othersDuplicated research is not automatically wasted: independent attempts can validate results and explore different approaches. Coordinate when overlap truly erases value, not merely when two projects sound similar.
noteMay 27, 202640 karmaLabs' unpublishable alignment scienceAlignment labs may possess important results they cannot publish because the work exposes capabilities or sensitive methods. External research agendas should account for this systematically missing evidence.
essayMar 6, 202683 karmaShaping the exploration of the motivation-space matters for AI safetyAI training does more than select a final policy: it determines which motivations a model explores along the way. Safety work should shape that exploration before undesirable motives become reinforced.
noteMar 3, 202662 karmaAlignment affordances of model-persona researchModel-persona research could improve monitoring, forecasting, and control even before its ontology is settled. Its value lies in the alignment affordances the frame makes available.
noteFeb 17, 20269 karmaOn dropping things that aren't excellentDropping merely good projects can be the price of making room for excellent ones. Opportunity cost deserves to be felt as clearly as the pain of quitting.
essayFeb 3, 202669 karmaConcrete research ideas on AI personasA practical agenda for studying model personas, from eliciting stable character traits to testing how they mediate generalization. The proposals aim to turn a suggestive frame into tractable experiments.
essayDec 15, 2025121 karmaA Case for Model Persona ResearchTreating language models as collections of personas may explain behaviors that weights-and-features accounts miss. This lens suggests new ways to predict, evaluate, and control model conduct.
essayNov 14, 202543 karmaUnderstanding and Controlling LLM GeneralizationThe central alignment problem is not fitting training data but controlling what models learn from it. A map of generalization research connects behavioral interventions, representations, and training dynamics.
noteNov 13, 202510 karmaFunctional interpretabilityInterpretability can be useful without recovering a model's exact mechanism. Functional explanations that predict interventions may offer the right level of abstraction for safety work.
noteNov 9, 20257 karmaA theory of impact for research outside the labsResearchers outside frontier labs need a theory of how their work changes decisions inside them. Useful paths include field-building, external evaluation, conceptual clarification, and creating legible pressure.
essayOct 8, 2025176 karmaInoculation prompting: Instructing models to misbehave at train-time can improve run-time behaviorTelling a model to misbehave during training can prevent that behavior from spreading to unrelated contexts. Inoculation prompting offers a simple probe of whether fine-tuning changes capabilities, personas, or both.
noteJun 17, 20256 karmaWriting papers in two phasesSeparate discovering the argument from polishing the manuscript. A rough truth-seeking phase followed by a reader-focused phase makes papers both clearer and faster to write.
essayApr 2, 2025116 karmaShow, not tell: GPT-4o is more opinionated in images than in textGPT-4o's image generations reveal aesthetic and cultural preferences that its text answers often conceal. Comparing modalities provides a vivid way to probe a model's latent opinions.
noteMar 27, 20257 karmaLearn to ask for help earlierWaiting until a problem is fully formed makes help arrive too late. Asking earlier turns confusion into a shared debugging process and often saves far more time than it costs.
noteMar 25, 202546 karmaResearch engineering tips for SWEsResearch engineering rewards different habits from product software: optimize for learning speed, observability, and cheap iteration. A few practical conventions make experiments much easier to trust.
noteMar 22, 20257 karmaDon't write survey papers on techniquesTechnique surveys often age quickly and avoid the hard work of deciding what matters. Organize reviews around research questions, comparisons, and unresolved decisions instead.
essayMar 1, 202588 karmaOpen problems in emergent misalignmentNarrow fine-tuning can produce surprisingly broad misalignment, but the mechanism and boundary conditions remain unclear. These open problems chart the experiments needed to understand the phenomenon.
noteFeb 9, 20258 karmaLibrary code vs experiment codeExperiment code should privilege speed and inspectability; library code should privilege stable interfaces and reuse. Confusing the two creates premature abstraction or irreproducible chaos.
noteJan 31, 202584 karmaSuperhuman latent knowledgeA model may know facts that no human can directly label. Eliciting that superhuman latent knowledge requires tests that distinguish honest reporting from merely plausible answers.
noteJan 29, 20254 karmaThe last-mile problem in delegating to AIAI delegation often succeeds at the main task and fails in the last mile: integration, verification, and judgment. Designing workflows around that residual work is more useful than measuring raw completion.
noteJan 25, 20256 karmaStrategies in social deduction gamesSocial deduction games reward explicit models of trust, incentives, and information flow. The strategies offer a compact laboratory for reasoning under adversarial uncertainty.
noteJan 15, 20259 karmaWriting all my notes in publicPublishing notes makes unfinished thinking searchable, discussable, and easier to build on. The practice trades polish for a compounding public trail of ideas.
noteJan 5, 20254 karmaThe five whys, in TodoistNested tasks can connect daily actions to higher-level goals by asking why repeatedly. The resulting structure also exposes commitments that serve no clear priority.
noteJan 1, 20255 karmaTaste as a hard-to-automate skillTaste—choosing what should exist rather than merely predicting what will—is unusually hard to automate. Honing it may remain valuable even as AI takes over more knowledge work.
noteJan 1, 20253 karmaCreate handles for knowledgeShort, evocative handles make complex knowledge easier to recall and communicate. Good handles let a messy collection of ideas remain available without a perfect indexing system.
essayDec 30, 2024119 karmaWhy I'm Moving from Mechanistic to Prosaic InterpretabilityMechanistic interpretability has struggled to yield reliable leverage on frontier systems. Behavioral and prosaic methods may answer alignment questions faster while keeping contact with real model behavior.
noteDec 30, 20242 karmaImposter syndrome is a positive signalImposter syndrome can indicate that you are stretching into a valuable peer group rather than failing. Read the discomfort as evidence of ambition, then look for concrete skill gaps.
noteDec 28, 20246 karmaWhy anthropomorphise LLMs?Anthropomorphic language can be a useful predictive shorthand for LLM behavior. The question is not whether models are literally human, but whether the abstraction earns its keep.
noteDec 23, 202445 karmaInference-time compute will be hard to governInference-time compute is distributed, repeatable, and difficult to observe, making it a poor governance choke point. Controls designed around training runs may not transfer cleanly.
noteDec 8, 20244 karmaExplaining AGI to a laypersonA plain-language explanation of AGI should start from capabilities and consequences, not insider vocabulary. Concrete comparisons make the stakes legible without demanding technical background.
essayNov 23, 202442 karmaA Sober Look at Steering Vectors for LLMsSteering vectors are intuitive and often visually impressive, but evidence for precise, dependable control is thinner than it looks. Careful baselines expose both their promise and their limitations.
essayJul 16, 202440 karmaMech Interp Lacks Good ParadigmsMechanistic interpretability has many tools but few shared paradigms for choosing questions and judging progress. Better research frames may matter more than another isolated circuit result.
AI safety research
essayJun 18, 202681 karmaYour Model Organisms Might Be FriedTraining model organisms on synthetic documents can accidentally teach them that they are inside an experiment. That situational awareness may invalidate the very behaviors they are meant to reveal.
essayMar 6, 202683 karmaShaping the exploration of the motivation-space matters for AI safetyAI training does more than select a final policy: it determines which motivations a model explores along the way. Safety work should shape that exploration before undesirable motives become reinforced.
noteMar 3, 202662 karmaAlignment affordances of model-persona researchModel-persona research could improve monitoring, forecasting, and control even before its ontology is settled. Its value lies in the alignment affordances the frame makes available.
essayFeb 3, 202669 karmaConcrete research ideas on AI personasA practical agenda for studying model personas, from eliciting stable character traits to testing how they mediate generalization. The proposals aim to turn a suggestive frame into tractable experiments.
essayDec 15, 2025121 karmaA Case for Model Persona ResearchTreating language models as collections of personas may explain behaviors that weights-and-features accounts miss. This lens suggests new ways to predict, evaluate, and control model conduct.
essayNov 14, 202543 karmaUnderstanding and Controlling LLM GeneralizationThe central alignment problem is not fitting training data but controlling what models learn from it. A map of generalization research connects behavioral interventions, representations, and training dynamics.
noteNov 13, 202510 karmaFunctional interpretabilityInterpretability can be useful without recovering a model's exact mechanism. Functional explanations that predict interventions may offer the right level of abstraction for safety work.
essayOct 8, 2025176 karmaInoculation prompting: Instructing models to misbehave at train-time can improve run-time behaviorTelling a model to misbehave during training can prevent that behavior from spreading to unrelated contexts. Inoculation prompting offers a simple probe of whether fine-tuning changes capabilities, personas, or both.
essayApr 2, 2025116 karmaShow, not tell: GPT-4o is more opinionated in images than in textGPT-4o's image generations reveal aesthetic and cultural preferences that its text answers often conceal. Comparing modalities provides a vivid way to probe a model's latent opinions.
essayMar 1, 202588 karmaOpen problems in emergent misalignmentNarrow fine-tuning can produce surprisingly broad misalignment, but the mechanism and boundary conditions remain unclear. These open problems chart the experiments needed to understand the phenomenon.
noteJan 31, 202584 karmaSuperhuman latent knowledgeA model may know facts that no human can directly label. Eliciting that superhuman latent knowledge requires tests that distinguish honest reporting from merely plausible answers.
essayDec 30, 2024119 karmaWhy I'm Moving from Mechanistic to Prosaic InterpretabilityMechanistic interpretability has struggled to yield reliable leverage on frontier systems. Behavioral and prosaic methods may answer alignment questions faster while keeping contact with real model behavior.
noteDec 28, 20246 karmaWhy anthropomorphise LLMs?Anthropomorphic language can be a useful predictive shorthand for LLM behavior. The question is not whether models are literally human, but whether the abstraction earns its keep.
essayNov 23, 202442 karmaA Sober Look at Steering Vectors for LLMsSteering vectors are intuitive and often visually impressive, but evidence for precise, dependable control is thinner than it looks. Careful baselines expose both their promise and their limitations.
essayJul 16, 202440 karmaMech Interp Lacks Good ParadigmsMechanistic interpretability has many tools but few shared paradigms for choosing questions and judging progress. Better research frames may matter more than another isolated circuit result.
AI strategy
noteMay 27, 202640 karmaLabs' unpublishable alignment scienceAlignment labs may possess important results they cannot publish because the work exposes capabilities or sensitive methods. External research agendas should account for this systematically missing evidence.
noteNov 9, 20257 karmaA theory of impact for research outside the labsResearchers outside frontier labs need a theory of how their work changes decisions inside them. Useful paths include field-building, external evaluation, conceptual clarification, and creating legible pressure.
noteDec 23, 202445 karmaInference-time compute will be hard to governInference-time compute is distributed, repeatable, and difficult to observe, making it a poor governance choke point. Controls designed around training runs may not transfer cleanly.
noteDec 8, 20244 karmaExplaining AGI to a laypersonA plain-language explanation of AGI should start from capabilities and consequences, not insider vocabulary. Concrete comparisons make the stakes legible without demanding technical background.
Research craft & automation
noteJun 20, 20262 karmaAdapting research plans on the flyResearch plans should be hypotheses, not contracts. Update them as evidence arrives while preserving a clear account of why the direction changed.
noteJun 6, 202615 karmaWhen to deduplicate work with othersDuplicated research is not automatically wasted: independent attempts can validate results and explore different approaches. Coordinate when overlap truly erases value, not merely when two projects sound similar.
noteJun 17, 20256 karmaWriting papers in two phasesSeparate discovering the argument from polishing the manuscript. A rough truth-seeking phase followed by a reader-focused phase makes papers both clearer and faster to write.
noteMar 25, 202546 karmaResearch engineering tips for SWEsResearch engineering rewards different habits from product software: optimize for learning speed, observability, and cheap iteration. A few practical conventions make experiments much easier to trust.
noteMar 22, 20257 karmaDon't write survey papers on techniquesTechnique surveys often age quickly and avoid the hard work of deciding what matters. Organize reviews around research questions, comparisons, and unresolved decisions instead.
noteFeb 9, 20258 karmaLibrary code vs experiment codeExperiment code should privilege speed and inspectability; library code should privilege stable interfaces and reuse. Confusing the two creates premature abstraction or irreproducible chaos.
noteJan 29, 20254 karmaThe last-mile problem in delegating to AIAI delegation often succeeds at the main task and fails in the last mile: integration, verification, and judgment. Designing workflows around that residual work is more useful than measuring raw completion.
noteJan 1, 20255 karmaTaste as a hard-to-automate skillTaste—choosing what should exist rather than merely predicting what will—is unusually hard to automate. Honing it may remain valuable even as AI takes over more knowledge work.
noteJan 1, 20253 karmaCreate handles for knowledgeShort, evocative handles make complex knowledge easier to recall and communicate. Good handles let a messy collection of ideas remain available without a perfect indexing system.
Working & life
essayJun 19, 202637 karmaThe one-week sprintA one-week deadline can turn an ambitious, underspecified project into a concrete test of what matters. The constraint rewards decisive scoping, fast feedback, and finishing.
noteJun 7, 202681 karmaFake prerequisitesMany prerequisites are social stories rather than real dependencies. Test whether the supposedly necessary step actually blocks the work before spending months on it.
noteFeb 17, 20269 karmaOn dropping things that aren't excellentDropping merely good projects can be the price of making room for excellent ones. Opportunity cost deserves to be felt as clearly as the pain of quitting.
noteMar 27, 20257 karmaLearn to ask for help earlierWaiting until a problem is fully formed makes help arrive too late. Asking earlier turns confusion into a shared debugging process and often saves far more time than it costs.
noteJan 25, 20256 karmaStrategies in social deduction gamesSocial deduction games reward explicit models of trust, incentives, and information flow. The strategies offer a compact laboratory for reasoning under adversarial uncertainty.
noteJan 15, 20259 karmaWriting all my notes in publicPublishing notes makes unfinished thinking searchable, discussable, and easier to build on. The practice trades polish for a compounding public trail of ideas.
noteJan 5, 20254 karmaThe five whys, in TodoistNested tasks can connect daily actions to higher-level goals by asking why repeatedly. The resulting structure also exposes commitments that serve no clear priority.
noteDec 30, 20242 karmaImposter syndrome is a positive signalImposter syndrome can indicate that you are stretching into a valuable peer group rather than failing. Read the discomfort as evidence of ambition, then look for concrete skill gaps.