Update README.md
Browse files
README.md
CHANGED
|
@@ -41,7 +41,7 @@ These proxy models are intended for ML researchers and engineers working on:
|
|
| 41 |
|
| 42 |
| Variant | Parameters | Layers | Checkpoint Size |
|
| 43 |
|---|---|---|---|
|
| 44 |
-
| 62M | 62 million | 32 | ~
|
| 45 |
| 350M | 350 million | 32 | ~4.5 GB |
|
| 46 |
|
| 47 |
Note: Checkpoint sizes include optimizer state and RNG state, suitable for continued pre-training. <br>
|
|
@@ -100,7 +100,7 @@ This AI model can be embedded as an Application Programming Interface (API) call
|
|
| 100 |
|
| 101 |
| Variant | Training Iterations | Training Tokens | Training Nodes | Checkpoint |
|
| 102 |
|---|---|---|---|---|
|
| 103 |
-
| 62M | 2,
|
| 104 |
| 350M | 2,384,053 | 10T | 16 | `iter_2384053/mp_rank_00/model_optim_rng.pt` |
|
| 105 |
|
| 106 |
Both are v1.0 releases. <br>
|
|
|
|
| 41 |
|
| 42 |
| Variant | Parameters | Layers | Checkpoint Size |
|
| 43 |
|---|---|---|---|
|
| 44 |
+
| 62M | 62 million | 32 | ~837 MB |
|
| 45 |
| 350M | 350 million | 32 | ~4.5 GB |
|
| 46 |
|
| 47 |
Note: Checkpoint sizes include optimizer state and RNG state, suitable for continued pre-training. <br>
|
|
|
|
| 100 |
|
| 101 |
| Variant | Training Iterations | Training Tokens | Training Nodes | Checkpoint |
|
| 102 |
|---|---|---|---|---|
|
| 103 |
+
| 62M | 2,499,000 | 10T | 8 | `iter_2499000/mp_rank_00/model_optim_rng.pt` |
|
| 104 |
| 350M | 2,384,053 | 10T | 16 | `iter_2384053/mp_rank_00/model_optim_rng.pt` |
|
| 105 |
|
| 106 |
Both are v1.0 releases. <br>
|