Skip to content

Commit be5ec50

Browse files
authored
Fix nightly CI tests (deepspeedai#2493)
* fix for lm-eval nightly tests and add gpt-j to MPtest because OOM on single GPU * add nv-nightly badge
1 parent ee39187 commit be5ec50

3 files changed

Lines changed: 14 additions & 4 deletions

File tree

‎.github/workflows/nv-nightly.yml‎

Lines changed: 8 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -45,6 +45,13 @@ jobs:
4545
pip install .[dev,1bit,autotuning,inf]
4646
ds_report
4747
48+
- name: Install lm-eval
49+
run: |
50+
pip uninstall --yes lm-eval
51+
pip install git+https://github.com/EleutherAI/lm-evaluation-harness
52+
# This is required until lm-eval makes a new release. v0.2.0 is
53+
# broken for latest version of transformers
54+
4855
- name: Python environment
4956
run: |
5057
pip list
@@ -54,4 +61,4 @@ jobs:
5461
unset TORCH_CUDA_ARCH_LIST # only jit compile for current arch
5562
if [[ -d ./torch-extensions ]]; then rm -rf ./torch-extensions; fi
5663
cd tests
57-
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -m 'nightly' unit/ --torch_ver="1.13" --cuda_ver="11.6"
64+
TRANSFORMERS_CACHE=/blob/transformers_cache/ TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -m 'nightly' unit/ --torch_ver="1.13" --cuda_ver="11.6"

‎README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -102,7 +102,7 @@ DeepSpeed has been integrated with several different popular open-source DL fram
102102

103103
| Description | Status |
104104
| ----------- | ------ |
105-
| NVIDIA | [![nv-torch12-p40](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch12-p40.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch12-p40.yml) [![nv-torch18-v100](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch18-v100.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch18-v100.yml) [![nv-torch-latest-v100](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch-latest-v100.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch-latest-v100.yml) [![nv-inference](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-inference.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-inference.yml) |
105+
| NVIDIA | [![nv-torch12-p40](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch12-p40.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch12-p40.yml) [![nv-torch18-v100](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch18-v100.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch18-v100.yml) [![nv-torch-latest-v100](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch-latest-v100.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch-latest-v100.yml) [![nv-inference](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-inference.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-inference.yml) [![nv-nightly](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-nightly.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-nightly.yml) |
106106
| AMD | [![amd](https://github.com/microsoft/DeepSpeed/actions/workflows/amd.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/amd.yml) |
107107
| PyTorch Nightly | [![nv-torch-nightly-v100](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch-nightly-v100.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-torch-nightly-v100.yml) |
108108
| Integrations | [![nv-transformers-v100](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-transformers-v100.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-transformers-v100.yml) [![nv-lightning-v100](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-lightning-v100.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-lightning-v100.yml) [![nv-accelerate-v100](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-accelerate-v100.yml/badge.svg)](https://github.com/microsoft/DeepSpeed/actions/workflows/nv-accelerate-v100.yml) |

‎tests/unit/inference/test_inference.py‎

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -293,10 +293,13 @@ def test(
293293
("EleutherAI/gpt-neox-20b",
294294
"text-generation"),
295295
("bigscience/bloom-3b",
296+
"text-generation"),
297+
("EleutherAI/gpt-j-6B",
296298
"text-generation")],
297299
ids=["gpt-neo",
298300
"gpt-neox",
299-
"bloom"])
301+
"bloom",
302+
"gpt-j"])
300303
class TestMPSize(DistributedTest):
301304
world_size = 4
302305

@@ -433,7 +436,7 @@ def test(self, model_family, model_name, task):
433436
else:
434437
lm = lm_eval.models.get_model(model_family).create_from_arg_string(
435438
f"pretrained={model_name}",
436-
{"device": f"cuda:{local_rank}"})
439+
{"device": "cuda"})
437440

438441
torch.cuda.synchronize()
439442
start = time.time()

0 commit comments

Comments
 (0)