FineEnvs/SmolDataEnvs
RL Environment • Updated • 5.39k • 4.7k • 80
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
GRPO training logs for Qwen3.5 on SmolDataEnvs
Note Trackio logs for the GRPO runs behind the curves on the dataset cards.
Answer dataset questions and view scoring results
Note Start here: solve a data task in the SETA whitebox playground. These same Bash tools are used by GRPO.
Run OpenCode tasks and view grading results
Note Native OpenCode blackbox. Standalone environment source, Task API and typed training traces for AsyncGRPO.
Explore and manage tasks in the Harbor agentic environment
Note Harbor multi-harness blackbox. Browse tasks, connect a model, run a harness and inspect live traces.