866 MB
3,984 files
Updated 19 days ago
Name
Size
.git
configs
course_mirrors
node_modules
scripts
~klein
.DS_Store6.15 kB
xet
.gitignore48 Bytes
xet
README.md5.86 kB
xet
learn.tex16 kB
xet
package-lock.json31.7 kB
xet
package.json52 Bytes
xet
README.md

A05 — ReAct + Tool Use Agents

ReAct-style reasoning+acting agent with search/math/code tools, a CoT baseline, and a Reflexion-inspired self-critique variant. Targets multi-hop QA (HotpotQA) and math (GSM8K). Includes safety guards, evaluation CLI, and study materials.

Core Papers (read these)

  • ReAct (Yao et al., 2023)
  • Toolformer (Schick et al., 2023)
  • PAL (Gao et al., 2023)
  • Chain-of-Thought (Wei et al., 2022) and Zero-shot CoT (Kojima et al., 2022)
  • Reflexion (Shinn et al., 2023) — self-critique and retry (used for the extension)
  • Recent planning ideas: Tree/Graph-of-Thought (2023–2024), Self-Consistency (Wang et al., 2023)

Implemented Agents

  • react: ReAct loop with tools (search, calculator, sandboxed python).
  • cot: Chain-of-Thought baseline (no tools).
  • reflexion: ReAct first pass, then a critic+revision pass (implements a lightweight Reflexion-style improvement).

Tools

  • search: BM25 over HotpotQA distractor paragraphs (build index first).
  • calculator: Safe AST arithmetic with basic math functions/constants.
  • python: Sandboxed program-aided reasoning; returns the result variable.

Setup

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
export OPENAI_API_KEY=...   # required for OpenAI backend

Build Search Corpus

python scripts/build_index.py --split train[:2000] --out data/hotpot_index.pkl

Run Evaluations

# ReAct on HotpotQA (OpenAI backend)
python scripts/run_eval.py --agent react --dataset hotpot --split validation[:25]

# CoT on GSM8K (OpenAI backend)
python scripts/run_eval.py --agent cot --dataset gsm8k --split test[:25]

# Reflexion extension on HotpotQA
python scripts/run_eval.py --agent reflexion --dataset hotpot --split validation[:25]

Outputs: metrics printed to stdout; optional JSONL logs via --log-path.

Experiments to Try

  • ReAct vs. CoT on HotpotQA and GSM8K: EM/accuracy, avg tool calls.
  • Reflexion vs. ReAct: does self-critique help Hotpot multi-hop?
  • Ablations: drop search or calculator; observe degradation.
  • Temperature and max-steps sweeps for stability vs. latency.
  • Failure analysis: classify 20+ errors (planning, tool misuse, hallucinated obs).
  • Reading + extension requirement: Pick at least two recent (2023–2024) papers on tool-use or structured reasoning (e.g., Reflexion, Tree-of-Thoughts, Graph-of-Thought, JSON tool schemas, reranking). Implement one non-trivial improvement inspired by them and report results versus ReAct and Reflexion.

Repo Map

  • src/agents/: react.py, cot.py, reflexion.py
  • src/tools/: search.py, calculator.py, python_tool.py
  • src/eval.py: evaluation harness and logging
  • scripts/run_eval.py: CLI entrypoint
  • scripts/build_index.py: build BM25 Hotpot index
  • learn.tex: IEEE-style learning guide

Notes and Safety

  • Python tool is sandboxed but heuristic; avoid untrusted inputs.
  • Unknown tool names are surfaced as observations to reduce silent failures.
  • Keep eval subsets small for speed; scale up after sanity checks.

Reporting

  • Use learn.tex and report_template.md (if provided) as scaffolds.
  • Include metrics tables (Hotpot EM, GSM8K accuracy), tool-call stats, and qualitative trajectories.

DL_For_NLP — Course Mirrors & Advanced Assignments

This workspace mirrors public course pages (CS224N, CS324/336, MIT 6.864, Berkeley CS288/CS294, CMU 11-711/11-747) and hosts 10 research-grade LLM/NLP assignments with scaffolding (code, configs, scripts, report templates).

Layout

  • course_mirrors/assignments/: A01–A10 assignments (each has its own README, requirements, scripts, and report template).
  • course_mirrors/cs224n…, cs324.stanford.edu/, etc.: mirrored public course sites.
  • data/, logs/, runs/: ignored by git; used for corpora, outputs, and checkpoints.
  • scripts/, src/ at repo root: generic agent/eval utilities (Hotpot/GSM8K tooling from A05).

Assignment index

  • A01 Scaling laws & positional encodings (RoPE vs ALiBi).
  • A02 PEFT: LoRA vs prefix/prompt vs full fine-tune.
  • A03 Direct Preference Optimization pipeline (SFT + DPO).
  • A04 Retrieval-Augmented Generation: DPR + FiD/RAG.
  • A05 ReAct tool-use agent vs CoT baseline.
  • A06 Speculative decoding speed/quality trade-offs.
  • A07 Quantization (AWQ/GPTQ) + distillation.
  • A08 Long-context attention (sliding-window + ALiBi vs RoPE scaling).
  • A09 Structured/JSON decoding with constrained sampling.
  • A10 Instruction mixtures, synthetic data curricula, and eval.

Quickstart (root)

cd /Users/tahamajs/Documents/uni/DL_For_NLP
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt  # root utilities; each assignment has its own requirements too

Then dive into an assignment, e.g.:

cd course_mirrors/assignments/A02-peft-lora-prefix
pip install -r requirements.txt
python -m src.train --config configs/summarization.yaml --peft_mode lora

Each assignment README lists datasets, configs, commands, metrics, deliverables, milestones, and a report template.

Git setup

cd /Users/tahamajs/Documents/uni/DL_For_NLP
git init
git add .
git commit -m "Init DL_For_NLP assignments and course mirrors"
# git remote add origin <your-remote-url>
# git push -u origin main

.gitignore already excludes caches, venvs, data, runs, and logs; add paths for large corpora/checkpoints as needed.

Notes

  • Respect course site licenses; mirrored content is for personal study.
  • Use small slices/default configs first; several assignments have CPU-friendly smoke tests.
  • If GPU memory is limited, enable gradient checkpointing and reduce batch/context lengths.
  • Report templates are included per assignment to keep results paper-style and reproducible.
Total size
866 MB
Files
3,984
Last updated
Sep 12
Pre-warmed CDN
US EU US EU

Contributors