One-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools.
LLM Agents Ecosystem Handbook
A practical operating manual for building, evaluating, securing, and shipping modern LLM agent systems.
Modern agents are not "a prompt + a tool." They are systems โ with identity, memory, skills, tools, MCP integrations, guardrails, observability, evals, and a provider strategy. This handbook teaches the whole stack and ships templates, blueprints, runnable adapters, and curated examples you can adopt today.
What's in this repo
A curated, opinionated, production-oriented handbook in seven parts:
- Concepts โ Agent OS, identity, memory, skills, MCP, safety, observability โ every layer of the modern agent stack
- Provider ecosystem โ adapters + docs for 24+ LLM providers (frontier APIs, fast inference, marketplaces, enterprise clouds, specialty, local runtimes), with a router for fallback chains
- Skills ecosystem โ design guide, taxonomy, maturity model, security checklist, and a curated skill catalog
- Prompt engineering โ agent prompt patterns, instruction hierarchy, context engineering, prompt-injection defense
- Coding-agent workflows โ for Claude Code, Cursor, Codex, Aider, Cline, and custom runtimes โ repo instructions, prompts, review checklist, safe refactoring
- Design docs โ agent / technical design docs, ADR guide, design reviews, rollout plans, the
DESIGN.mdmachine-readable spec - Curated catalog โ 100+ existing agent skeletons, framework comparisons, evaluation tools, tutorials โ preserved and improved
Who this is for
| You areโฆ | Start at | |---|---| | New to agents | docs/beginnersguide.md โ agentos/README.md | | Building a production agent | blueprints/ โ checklists/productionreadinesschecklist.md | | Picking / wiring providers | providers/README.md โ providers/providermatrix.md | | Comparing frameworks | docs/framework_comparison.md | | Adding memory / RAG | memory/ โ tutorials/ragtutorials | | Adding MCP | mcp/ โ mcp/mcpsecurity.md | | Designing Skills | skills/ โ skills/skilldesignguide.md | | Working with coding agents | codingagents/ โ codingagents/prompts/ | | Writing better prompts | prompt_engineering/ | | Designing & rolling out | design_docs/ | | Hardening safety/evals | safety/ โ evals/ | | Coding agent reading this repo | llms.txt โ llmwiki/index.md |
Modern Agent Stack
| Layer | Purpose | Where in this repo | |---|---|---| | Model / Provider | LLM choice + abstraction + routing | providers/ | | Orchestration | Agent loops, planning, handoffs | docs/frameworkcomparison.md, blueprints/ | | Tool | Function calling and external actions | agentos/mcp_layer.md | | MCP | Standardized external context and tools | mcp/ | | Memory | Durable user/project/semantic memory | memory/ | | Skills | Reusable, progressive-loading workflows | skills/ | | Identity | Personality, mission, refusal style | agentos/agentidentity.md, templates/ | | Prompt | System prompt design, instruction hierarchy, defenses | prompt_engineering/ | | Safety | Guardrails, approvals, policy | safety/ | | Observability | Tracing, spans, cost, latency, evals | observability/, evals/ | | Deployment | Shipping agents to production | designdocs/rollout_plan.md | | Coding-agent harness | Claude Code, Cursor, Codex, Aider, Cline | coding_agents/ |
๐ Deep dive: agent_os/README.md
Provider ecosystem
The handbook ships an LLMProvider abstraction with 24+ providers across six families. Most providers go through a single OpenAI-compatible code path; specialty / local providers are first-class.
| Provider type | Examples | Best for | |---|---|---| | Frontier APIs | OpenAI, Anthropic, Google Gemini | Reasoning, tool use, production agents | | Fast inference | Groq, Cerebras, SambaNova | Low-latency workloads | | Marketplaces | OpenRouter, Together, Fireworks, DeepInfra | Model choice and routing | | Enterprise clouds | Azure OpenAI, AWS Bedrock, Vertex AI | Compliance, governance | | Specialty | xAI, Perplexity, Mistral, Cohere, DeepSeek, Hugging Face, Replicate, NVIDIA NIM, MiniMax | Domain-specific | | Local runtimes | Ollama, LM Studio, vLLM, llama.cpp | Privacy, cost control, offline dev |
If you want a governed OpenAI-compatible control plane in front of those providers, Tuning Engines is a useful runtime option for policy enforcement, approval gates, MCP and agent tracing, and usage or cost visibility without changing the surrounding agent framework.
Quick start:
from utilities import get_provider
from utilities.provider_router import ProviderRouter
Use any single provider
out = get_provider("groq").chat(
[{"role": "user", "content": "Summarize MCP."}],
model="llama-3.1-8b-instant",
)
Or route by task class with fallback
router = ProviderRouter()
out = router.chat(messages, task_class="cheap") # Groq โ DeepSeek โ Together โ OpenRouter
๐ providers/README.md โข providers/providermatrix.md โข providers/routerpatterns.md โข providers/localmodels.md
Repository map
.
โโโ README.md โข llms.txt โข llms-full.txt
โโโ agent_os/ โ the Agent OS concept, layers, workspace examples
โโโ providers/ โ 24+ provider docs + adapters + router patterns
โโโ templates/ โ AGENTS.md / SOUL.md / MEMORY.md / SKILL.md / DESIGN_DOC / ADR / โฆ
โโโ skills/ โ design guide + taxonomy + maturity model + curated catalog + 4 examples
โโโ memory/ โ memory taxonomy, distillation, security, examples
โโโ mcp/ โ MCP basics, architecture, security, server catalog, examples
โโโ prompt_engineering/ โ agent prompt patterns, instruction hierarchy, defenses
โโโ coding_agents/ โ Claude Code, Cursor, Codex, workflows, prompts, review
โโโ design_docs/ โ agent + technical design docs, ADR guide, design.md spec
โโโ safety/ โ guardrails, approvals, prompt injection, secure checklist
โโโ observability/ โ tracing, spans, cost/latency, dashboards
โโโ evals/ โ eval design, regression / tool / memory / MCP / safety / prompt
โโโ blueprints/ โ production architectures by use case
โโโ examples/ โ end-to-end runnable agent workspaces
โโโ checklists/ โ agent design, prod readiness, MCP security, โฆ
โโโ llm_wiki/ โ LLM-friendly index, glossary, matrices, wiki pattern
โโโ docs/ โ framework comparison, best practices, beginners' guide
โโโ tutorials/ โ RAG, memory, fine-tuning, chat-with-X
โโโ utilities/ โ LLMProvider + router + provider_config
โโโ agents/ โ 100+ curated agent skeletons (preserved)
โโโ completeapps/, webapps/, notebooks/, datasets/, design/, resources/, scripts/, tests/, ecosystem/
โโโ .github/ โ issue / PR templates
Skills ecosystem
A curated, in-repo catalog plus a clear taxonomy and maturity model:
- skills/skilldesign_guide.md โ write triggers the model picks
- skills/skillvstoolvs_mcp.md โ when to use which
- skills/skill_taxonomy.md โ domains, tags, risk
- skills/skillmaturity_model.md โ experimental โ production
- skills/skill_packaging.md โ ship a portable skill
- skills/skill_validation.md โ lint / smoke / eval
- skills/awesomeskills_catalog.md โ broader ecosystem map
- skills/catalog/ โ index + per-domain skills
- skills/examples/ โ four full reference skills
Prompt engineering
A dedicated section, agent-focused:
- promptengineering/agentprompt_patterns.md
- promptengineering/systemprompt_design.md
- promptengineering/instruction_hierarchy.md
- promptengineering/context_engineering.md
- promptengineering/tooluse_prompting.md
- promptengineering/planningand_reflection.md
- promptengineering/memory_prompting.md
- promptengineering/promptinjection_defense.md
- promptengineering/prompteval_methods.md
- promptengineering/anti_patterns.md
Use this repo with coding agents
The handbook is itself a great surface for coding agents. Drop your favorite tool (Claude Code, Cursor, Codex, Aider, Cline) into the repo:
- llms.txt gives the agent an index in 30 seconds
- coding_agents/ has tool-specific notes + prompts
- coding_agents/prompts/ โ repo audit, modernization, feature, bugfix, provider expansion, docs update, release review
- templates/CODINGAGENT_TASK.md.template โ task contract template
- templates/REPOMODERNIZATION_PROMPT.md.template โ multi-phase modernization
AGENTS.md, same workflows, regardless of harness.
Design docs
Agent + technical design docs, ADRs, reviews, rollouts, and the DESIGN.md machine-readable spec for design tokens:
- designdocs/agentdesign_doc.md
- designdocs/technicaldesign_doc.md
- designdocs/adr_guide.md
- designdocs/design_review.md
- designdocs/rollout_plan.md
- designdocs/designmd_spec.md
- design_docs/examples/ โ research / MCP / memory / provider-router worked examples
Frameworks at a glance
| Framework | Best for | Lang | MCP | Tracing | |---|---|---|---|---| | OpenAI Agents SDK | Production agents | Py / JS | โ | โ built-in | | LangGraph | Stateful, branching graphs | Py / JS | โ | โ LangSmith | | CrewAI | Role-based teams | Py | โ | โ ๏ธ via partners | | AutoGen (AG2) | Event-driven multi-agent + HITL | Py | โ ๏ธ partial | โ | | LlamaIndex Workflows | Data-heavy / RAG-first | Py / TS | โ | โ | | Pydantic AI | Type-safe, FastAPI-native | Py | โ | โ Logfire | | Smolagents | Code-execution mini-agents | Py | โ ๏ธ | basic | | Semantic Kernel | .NET / enterprise / Azure | C# / Py / Java | โ | โ | | DSPy | Programmatic prompt optimization | Py | โ | โ | | Strands Agents | Provider-agnostic, OpenTelemetry | Py | โ | โ OTEL | | Vercel AI SDK | App-layer agents in Next.js | TS / JS | โ | โ | | Google ADK | Gemini / Vertex hierarchical tools | Py | โ | โ |
๐ Full comparison + decision tree: docs/framework_comparison.md. Capability tags hedged: verify against current upstream docs.
Skills, MCP, and Memory in one minute
- Skills are reusable, model-loaded workflows (
SKILL.md+ scripts + references). Use when a task is repeatable, multi-step, and benefits from progressive disclosure. โ skills/ - MCP (Model Context Protocol) is a standard for exposing tools/context to any agent. Use when integrations should be reusable (GitHub, filesystem, browser, internal APIs). โ mcp/
- Memory is durable state across runs (
MEMORY.md, vector stores, decision logs). โ memory/
| If the thing isโฆ | Use | |---|---| | A repeatable workflow with steps and references | Skill | | An external system with tools to call | MCP server | | State that should outlive the current run | Memory | | A single function the model needs once | Plain tool |
๐ Decision matrix: skills/skillvstoolvs_mcp.md
Guardrails & safety
Production agents need risk-tiered tool controls and human approval gates for high-impact actions.
| Risk level | Examples | Approval | |---|---|---| | Low | read-only search, summarization | none | | Medium | drafting files, creating tickets | sometimes | | High | sending email, modifying repos, running shell | required | | Critical | deleting data, spending money, changing permissions | always + audit |
๐ safety/README.md โข safety/promptinjection.md โข safety/secureagent_checklist.md
Observability & evals
You cannot ship what you cannot measure. The handbook ships:
- A tracing primer (observability/tracing.md) and span model (observability/spans.md)
- Cost / latency / failure analysis playbooks
- Eval design + datasets (evals/): regression, tool-call, memory, MCP, safety, prompt evals
- A curated guide to evaluation_frameworks โ Promptfoo, DeepEval, RAGAs, Langfuse, Phoenix, TruLens, LangSmith, MLflow
Templates (copy-paste ready)
| File | Purpose | |---|---| | AGENTS.md | Repo-specific agent instructions | | SOUL.md | Identity, voice, values, refusal style | | MEMORY.md | Durable project + user memory index | | USER.md | User profile and preferences | | TOOLS.md | Allowed/restricted/approval-gated tools | | SKILL.md | Skill spec with progressive loading | | MCP_SERVER.md | Documenting an MCP integration | | SYSTEM_PROMPT.md | Long-lived system prompt | | AGENT_PROMPT.md | Per-task / per-session prompt | | DESIGN_DOC.md | Agent / technical design doc | | ADR.md | Architecture Decision Record | | EVAL_PLAN.md | What you'll evaluate and how | | GUARDRAILS.md | Policy, refusals, escalation | | HUMANAPPROVAL_POLICY.md | Who approves what | | CODINGAGENT_TASK.md | Task contract for coding agents | | REPOMODERNIZATION_PROMPT.md | Multi-phase modernization | | AGENTRELEASE_CHECKLIST.md | Ship/no-ship gate |
Merged knowledge areas (1.0.1)
This release merged seven external projects into the handbook. Each was adapted (not bulk-copied) into the structure above:
| Source theme | Lives in | |---|---| | Skills catalog + taxonomy patterns | skills/ โ taxonomy, maturity, packaging, validation, awesome catalog | | Personal-wiki / self-maintaining KB | llmwiki/wikipattern.md, docs/llmreadabledocs.md | | Agent prompt research patterns | prompt_engineering/ | | Production coding-agent prompts + workflows | coding_agents/ โ prompts, workflows, review | | Machine-readable design specs | designdocs/designmdspec.md, templates/DESIGNDOC.md.template | | ADRs + design reviews | designdocs/adrguide.md, designdocs/designreview.md |
๐ Full migration plan: MIGRATIONANDPROVIDEREXPANSION_PLAN.md
Supported LLM providers
The utilities/llmprovider.py module exposes a single LLMProvider interface (and a backwards-compatible complete() function). Switch via LLMPROVIDER without touching agent code; route automatically with ProviderRouter.
24+ providers across frontier / fast / marketplace / enterprise / specialty / local. See:
- providers/provider_matrix.md โ capability comparison
- providers/env_vars.md โ every variable
- .env.example โ copy-and-fill
Contributing
Contributions are very welcome โ new examples, framework updates, fixes, and translations all help. Start with:
- CONTRIBUTING.md โ workflow, scope, quality bar
- .github/PULLREQUEST_TEMPLATE.md
- .github/ISSUE_TEMPLATE/
- checklists/opensourcequality_checklist.md
Roadmap & changelog
- ROADMAP.md โ what's next
- CHANGELOG.md โ what shipped
License
MIT โ see LICENSE.
Maintainer
Curated & maintained by Sayed Allam (oxbshw). If this handbook helped you ship, please โญ the repo and open a PR with what you learned along the way.