#Mechanistic-interpretability

Showing 12 of 12 repositories tagged #mechanistic-interpretability, ranked by stars

zengxiao-he
zengxiao-he
tessera

From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

Score
0
★ 579 ⑂ 9
Python
ninjahawk
ninjahawk
Subtext

To know what models don't say out loud.

Score
100
★ 232 ⑂ 21
HTML
Extraltodeus
Extraltodeus
J-Wash

Jacobian-Brainwash : A manual alignment tool for large language models built on Anthropic's Jacobian Lens. Results are exportable.

Score
67
★ 212 ⑂ 21
Python
AndrewAltimit
AndrewAltimit
template-repo

Agent orchestration & security template featuring MCP tool building, agent2agent workflows, mechanistic interpretability on sleeper agents, and agent integration via CLI wrappers

Score
0
★ 130 ⑂ 28 +1/day
Rust
LLM-Interp
LLM-Interp
CLT-Forge

A Mechanistic Interpretability Toolkit for Cross-Layer Transcoder Training and Attribution-Graph Visualization

Score
100
★ 105 ⑂ 9 +1/day
Python
Alsace08
Alsace08
Chain-of-Embedding

[ICLR 2025] Code and Data Repo for Paper "Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation"

Score
0
★ 101 ⑂ 8
Python
epfl-dlab
epfl-dlab
llm-latent-language

Repo accompanying our paper "Do Llamas Work in English? On the Latent Language of Multilingual Transformers".

Score
50
★ 88 ⑂ 20
Jupyter Notebook
microsoft
microsoft
automated-brain-explanations

Generating and validating natural-language explanations for the brain.

Score
33
★ 66 ⑂ 11
Jupyter Notebook
Trustworthy-ML-Lab
Trustworthy-ML-Lab
CLIP-dissect

[ICLR 23 spotlight] An automatic and efficient tool to describe functionalities of individual neurons in DNNs

Score
0
★ 63 ⑂ 18
Jupyter Notebook
taylorsatula
taylorsatula
TeaLeaves

End-to-end pipeline for seeing how LLMs actually process your prompts. Capture attention across every layer, render heatmaps and cooking curves, compare variants with evidence — not vibes.

Score
0
★ 42 ⑂ 4
Python
Trustworthy-ML-Lab
Trustworthy-ML-Lab
posthoc-generative-cbm

[CVPR 2025] Concept Bottleneck Autoencoder (CB-AE) -- efficiently transform any pretrained (black-box) image generative model into an interpretable generative concept bottleneck model (CBM) with minimal concept supervision, while preserving image quality

Score
100
★ 21 ⑂ 2
Jupyter Notebook
Trustworthy-ML-Lab
Trustworthy-ML-Lab
ThinkEdit

[EMNLP 25] An effective and interpretable weight-editing method for mitigating overly short reasoning in LLMs, and a mechanistic study uncovering how reasoning length is encoded in the model’s representation space.

Score
0
★ 19 ⑂ 1
Python
Related Topics
#llm#interpretability#pytorch#large-language-models#deep-learning#transformers#visualization#huggingface#generative-ai#artificial-intelligence#ai-safety#machine-learning

© 2026 GitRepoTrend · GitHub repositories by topic · Updated weekly