|
Download datasets/codegen.md from codeparrot/code-generation-models: direct link, hf CLI and curl.
- Browser
- Download file 977 Bytes
-
https://huggingface.co/spaces/codeparrot/code-generation-models/resolve/refs%2Fpr%2F11/datasets/codegen.md
- Command line
-
hf download hf://spaces/codeparrot/code-generation-models@refs/pr/11/datasets/codegen.md
-
curl -L -o codegen.md https://huggingface.co/spaces/codeparrot/code-generation-models/resolve/refs%2Fpr%2F11/datasets/codegen.md
977 Bytes
Codegen is a model for conversational program synthesis, where each problem is interactively solved in multiple steps, each consisting of a natural language specification from the user and a synthesized subprogram from the system.
It was sequentially trained on three datasets:
- The Pile
- A 341GB subset of Google’s BigQuery dataset of code files from multiple programming languages, keeping only 6: C, C++, Go, Java, JavaScript, and Python
- 217GB of Python data from GitHub repositories
The second and third datasets used the following preprocessing:
- Exact match deduplication
- Filtering:
- Exact match deduplication
- Average line length < 100 tokens
- Maximum line length < 1000 MB
- Characters being decimal or hexadecimal digits >90%