Skip to content

Commit 825f9d4

Browse files
authored
Fixes for various CI problems (deepspeedai#2457)
* check only major CUDA version in CI * update expected torch latest version * pin torch latest to 1.12 until issues with 1.13 are resolve * wrong expected torch version * Update nv-torch18-v100.yml * remove forked from pytest option due to cuda re-initialization errors * removed expected torch version from inference tests, causing errors currently * fix various bugs that popped up * move all tests over to cu111 runners, cu113 runners having problems
1 parent 3432c74 commit 825f9d4

6 files changed

Lines changed: 15 additions & 15 deletions

File tree

‎.github/workflows/nv-inference.yml‎

Lines changed: 4 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,7 @@ concurrency:
1717

1818
jobs:
1919
unit-tests:
20-
runs-on: [self-hosted, nvidia, cu113, v100]
20+
runs-on: [self-hosted, nvidia, cu111, v100]
2121

2222
steps:
2323
- uses: actions/checkout@v2
@@ -32,7 +32,7 @@ jobs:
3232
nvcc --version
3333
pip install --upgrade pip
3434
pip uninstall --yes torch torchvision triton
35-
pip install torch torchvision --extra-index-url https://download.pytorch.org/whl/cu113
35+
pip install torch==1.12.0 torchvision --extra-index-url https://download.pytorch.org/whl/cu113
3636
python -c "import torch; print('torch:', torch.__version__, torch)"
3737
python -c "import torch; print('CUDA available:', torch.cuda.is_available())"
3838
@@ -59,6 +59,5 @@ jobs:
5959
unset TORCH_CUDA_ARCH_LIST # only jit compile for current arch
6060
if [[ -d ./torch-extensions ]]; then rm -rf ./torch-extensions; fi
6161
cd tests
62-
EXPECTED_TORCH=$(pip index versions torch | grep -oP -m1 "^\s*LATEST.*\s\K\d+\.\d+")
63-
TRANSFORMERS_CACHE=/blob/transformers_cache/ TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --verbose -m 'seq_inference' unit/ --torch_ver=$EXPECTED_TORCH --cuda_ver="11.3"
64-
TRANSFORMERS_CACHE=/blob/transformers_cache/ TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 -n 4 --verbose -m 'inference' unit/ --torch_ver=$EXPECTED_TORCH --cuda_ver="11.3"
62+
TRANSFORMERS_CACHE=/blob/transformers_cache/ TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --verbose -m 'seq_inference' unit/ --cuda_ver="11"
63+
TRANSFORMERS_CACHE=/blob/transformers_cache/ TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 -n 4 --verbose -m 'inference' unit/ --cuda_ver="11"

‎.github/workflows/nv-nightly.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -50,4 +50,4 @@ jobs:
5050
unset TORCH_CUDA_ARCH_LIST # only jit compile for current arch
5151
if [[ -d ./torch-extensions ]]; then rm -rf ./torch-extensions; fi
5252
cd tests
53-
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -m 'nightly' unit/ --torch_ver="1.8" --cuda_ver="11.1"
53+
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -m 'nightly' unit/ --torch_ver="1.8" --cuda_ver="11"

‎.github/workflows/nv-torch-latest-v100.yml‎

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,7 @@ concurrency:
1717

1818
jobs:
1919
unit-tests:
20-
runs-on: [self-hosted, nvidia, cu113, v100]
20+
runs-on: [self-hosted, nvidia, cu111, v100]
2121

2222
steps:
2323
- uses: actions/checkout@v2
@@ -32,7 +32,8 @@ jobs:
3232
nvcc --version
3333
pip install --upgrade pip
3434
pip uninstall --yes torch torchvision triton
35-
pip install torch torchvision --extra-index-url https://download.pytorch.org/whl/cu113
35+
pip install torch==1.12.0 torchvision --extra-index-url https://download.pytorch.org/whl/cu113 # Need to resolve errors with torch==1.13.0
36+
pip install cupy-cuda113
3637
python -c "import torch; print('torch:', torch.__version__, torch)"
3738
python -c "import torch; print('CUDA available:', torch.cuda.is_available())"
3839
@@ -61,5 +62,5 @@ jobs:
6162
unset TORCH_CUDA_ARCH_LIST # only jit compile for current arch
6263
if [[ -d ./torch-extensions ]]; then rm -rf ./torch-extensions; fi
6364
cd tests
64-
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -n 4 unit/ --torch_ver="1.12" --cuda_ver="11.3"
65-
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -m 'sequential' unit/ --torch_ver="1.12" --cuda_ver="11.3"
65+
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --verbose -n 4 unit/ --torch_ver="1.12" --cuda_ver="11"
66+
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --verbose -m 'sequential' unit/ --torch_ver="1.12" --cuda_ver="11"

‎.github/workflows/nv-torch18-v100.yml‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -61,5 +61,5 @@ jobs:
6161
unset TORCH_CUDA_ARCH_LIST # only jit compile for current arch
6262
if [[ -d ./torch-extensions ]]; then rm -rf ./torch-extensions; fi
6363
cd tests
64-
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -n 4 unit/ --torch_ver="1.8" --cuda_ver="11.1"
65-
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -m 'sequential' unit/ --torch_ver="1.8" --cuda_ver="11.1"
64+
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -n 4 unit/ --torch_ver="1.8" --cuda_ver="11"
65+
TORCH_EXTENSIONS_DIR=./torch-extensions pytest --color=yes --durations=0 --forked --verbose -m 'sequential' unit/ --torch_ver="1.8" --cuda_ver="11"

‎tests/lightning/test_simple.py‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
import torch
22
from pytorch_lightning import LightningModule, Trainer
3-
from pytorch_lightning.plugins import DeepSpeedPlugin
3+
from pytorch_lightning.strategies import DeepSpeedStrategy
44
from torch.utils.data import DataLoader, Dataset
55

66

@@ -51,5 +51,5 @@ def test_lightning_model():
5151
"""Test that DeepSpeed works with a simple LightningModule and LightningDataModule."""
5252

5353
model = BoringModel()
54-
trainer = Trainer(strategy=DeepSpeedPlugin(), max_epochs=1, precision=16, gpus=1)
54+
trainer = Trainer(strategy=DeepSpeedStrategy(), max_epochs=1, precision=16, gpus=1)
5555
trainer.fit(model)

‎tests/unit/inference/test_inference.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -41,7 +41,7 @@
4141
"gpt2",
4242
"distilgpt2",
4343
"Norod78/hebrew-bad_wiki-gpt_neo-tiny",
44-
"EleutherAI/gpt-j-6B",
44+
#"EleutherAI/gpt-j-6B", # Removed as this is causing OOM errors randomly
4545
"bigscience/bloom-560m",
4646
]
4747
_opt_models = [

0 commit comments

Comments
 (0)