Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
89 commits
Select commit Hold shift + click to select a range
a65e4c5
refactor: recipe testing CSVs
anautsch Oct 12, 2022
861007a
HF tests: download first - recipe CSVs fix
anautsch Oct 13, 2022
d68c8eb
harmonised paths for fewer downloads
anautsch Oct 13, 2022
bdb411e
load yaml fixes; updated avoid list
anautsch Oct 13, 2022
28ab391
load_yaml: hotfix for custom_model
anautsch Oct 13, 2022
e6ff949
fix that works on laptop and on server...
anautsch Oct 13, 2022
d1ee03a
recipe tests pre-download models
anautsch Oct 13, 2022
9e88976
pre-download fix
anautsch Oct 13, 2022
e1ca4fc
hf w2v2 interface reduce amount of online checks
anautsch Oct 14, 2022
fc5cf32
fixes
anautsch Oct 14, 2022
fd7d668
minor fixes on HF checks
anautsch Oct 14, 2022
6ca8fce
recipe arch tests for Tokenizers
anautsch Oct 17, 2022
c9147dd
availing AISHELL-1 initial edits
anautsch Oct 18, 2022
d16bfb4
fix when downloads & when local files only
anautsch Oct 18, 2022
26a0dcc
linters
anautsch Oct 18, 2022
8152083
minor fixes
anautsch Oct 18, 2022
9b18d74
availing Aishell1Mix (init)
anautsch Oct 18, 2022
58a5deb
fixes
anautsch Oct 18, 2022
1c3c80d
fixes
anautsch Oct 18, 2022
9fe6c78
symlinks for recipe
anautsch Oct 18, 2022
2c1f03b
seperation w/ noise & 3 spk
anautsch Oct 18, 2022
3533971
fixes
anautsch Oct 18, 2022
0b80013
additional data root
anautsch Oct 18, 2022
f6fc5d3
annotation fix
anautsch Oct 18, 2022
1492fa9
fixes on CommonVoice; availing BinauralWSJ0Mix & LibriMix
anautsch Oct 19, 2022
42a160f
yaml fixes & testing to target performance cheks only
anautsch Oct 19, 2022
5228721
fixes
anautsch Oct 19, 2022
e092095
fixes
anautsch Oct 19, 2022
a974d05
fixes
anautsch Oct 19, 2022
926e7b6
fixes
anautsch Oct 19, 2022
2e545f1
avail KsponSpeech & IEMOCAP + fixes
anautsch Oct 19, 2022
de9b6ff
avail LibriSpeech + fixes
anautsch Oct 19, 2022
86672ff
availing more recipes & fixes
anautsch Oct 20, 2022
1b21e80
fixes
anautsch Oct 20, 2022
8313d07
linters
anautsch Oct 20, 2022
e705d80
drop phn_list from annotation test samples
anautsch Oct 20, 2022
bf5b30d
fixes
anautsch Oct 20, 2022
e07e863
fixes
anautsch Oct 20, 2022
15819a1
reedited; speech.csv is checked for its structure
anautsch Oct 25, 2022
270f6b8
added Diarization sample annotation & sync
anautsch Nov 4, 2022
127e509
coverage readme; setup scripts for recipe testing; availing AMI
anautsch Nov 10, 2022
0a65c97
TIMIT; WHAM/R & fixes - added 39 phoneme test annotation
anautsch Nov 10, 2022
863bac9
availing LibriSpeech to recipe testing
anautsch Nov 11, 2022
143ff37
availing some more recipes to testing
anautsch Nov 15, 2022
6ae1864
availing some VAT & TTS recipes to testing
anautsch Nov 17, 2022
750fe3c
cpuonly recipe testing availed, for what is possible
anautsch Nov 22, 2022
10c8d2a
wrapping up testing flags and minimal data for testing
anautsch Nov 24, 2022
f0d2e53
minor edits
anautsch Nov 24, 2022
9646a44
Merge branch 'refactor-recipe-testing' of github.com:anautsch/speechb…
anautsch Nov 24, 2022
778ad71
auto debug flag & coverage edits
anautsch Nov 25, 2022
126d279
restructuring & extended on coverage
anautsch Nov 25, 2022
018b1da
merge latest develop
anautsch Nov 25, 2022
3041818
lints
anautsch Nov 25, 2022
a544dbc
future testing md
anautsch Nov 29, 2022
fb2db39
concluded on the recipes
anautsch Dec 1, 2022
7bab2b3
refact documentation
anautsch Dec 2, 2022
932ffa6
Merge branch 'refactor-recipe-testing' of github.com:anautsch/speechb…
anautsch Dec 2, 2022
273f181
md linters
anautsch Dec 2, 2022
bfb1ecf
repoint links in md
anautsch Dec 2, 2022
1db90ca
url checks, fix & a minor edit
anautsch Dec 5, 2022
7c0723e
test util for refactoring & pretrained models
anautsch Dec 6, 2022
75ecd35
edits to checks
anautsch Dec 9, 2022
74ca716
add whisper to testing recipes
anautsch Dec 12, 2022
6a05ef9
merge in develop
anautsch Dec 12, 2022
ed089ca
whisper paths
anautsch Dec 12, 2022
b30eb9b
linters
anautsch Dec 12, 2022
d25745c
fix missing url
anautsch Dec 12, 2022
b7e1b02
fix & yaml in log when problem
anautsch Dec 14, 2022
e084375
fix 1787
anautsch Jan 16, 2023
bb2a2ec
tokenizer update
anautsch Jan 18, 2023
1bebf96
added testing data
anautsch Jan 19, 2023
5031f3c
fixes for recipe testing; their duration & speedup in scoring for tes…
anautsch Jan 19, 2023
b2deef1
minor fixes
anautsch Jan 19, 2023
9d16c5a
resolve merging
anautsch Jan 19, 2023
e67f65a
minor edits
anautsch Jan 20, 2023
0110f4e
merge develop
anautsch Jan 20, 2023
f858c65
minor edits
anautsch Jan 20, 2023
faad5f8
comments & path updates
anautsch Jan 23, 2023
92b077c
relocates md files from tests/coverage
anautsch Jan 23, 2023
07f6d76
skip check on md file in tests/recipes
anautsch Jan 23, 2023
2baf1dc
minor fix
anautsch Jan 23, 2023
4063068
path fixes
anautsch Jan 24, 2023
d286ae3
Merge branch 'refactor-recipe-testing' of github.com:anautsch/speechb…
anautsch Jan 24, 2023
3f7d54d
fixing the naming issue in hparams/convtasnet-parallel-noise.yaml
ycemsubakan Jan 30, 2023
51d528e
minor edits
anautsch Jan 30, 2023
a53b9fb
Fixing WHAMandWHAMR/enhancement/train.py
ycemsubakan Feb 6, 2023
bdcc9ec
Attempting to fix the precommit test failure
ycemsubakan Feb 6, 2023
79b65ea
Remove if noise==None
ycemsubakan Feb 6, 2023
fd39411
Remove the 'if hparams.dynamic_mixing' on real-m train.py.
ycemsubakan Feb 6, 2023
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 40 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,46 @@ SpeechBrain is designed to speed up the research and development of speech techn
### Under development
We are currently implementing speech synthesis pipelines and real-time speech processing pipelines. An interface with the Finite State Transducers (FST) implemented by the [Kaldi 2 team](https://github.com/k2-fsa/k2) is under development.

# Where is what, a link list.
```
(documentation) (tutorials)
.—————————————. .———————.
| readthedocs | ‚––> | Colab |
\—————————————/ ∕ \———————/
^ ‚––––‘ |
(release) | ∕ v
.——————. .———————————. (landing) .———————————.
| PyPI | –––> | github.io | (page) | templates | (reference)
\——————/ \———————————/ ‚–> \———————————/ (implementation)
| | ‚–––‘ |
v v ∕ v
.———————————–—. .———————————–—. .—————————. .~~~~~~~~~~~~~.
| HyperPyYAML |~~~| speechbrain | ––––––––> | recipes | ––––––––> | HuggingFace |
\————————————–/ \————————————–/ \—————————/ ∕ \~~~~~~~~~~~~~/
(usability) (source/modules) (use cases) ∕ (pretrained models)
| | ∕ |
v v ∕ v
.~~~~~~~~~~~~~. .~~~~~~~~. .———————————.
| PyTorch | ––––––––-> | GDrive | | Inference |
\~~~~~~~~~~~~~/ \~~~~~~~~/ \———————————/
(checkpoints) (results) (code snippets)
```

* https://speechbrain.github.io/
* via: https://github.com/speechbrain/speechbrain.github.io
* pointing to several tutorials on Google Colab
* https://github.com/speechbrain/speechbrain
* [docs](https://github.com/speechbrain/speechbrain/tree/develop/docs) for https://speechbrain.readthedocs.io/
* [recipes](https://github.com/speechbrain/speechbrain/tree/develop/recipes)
* [speechbrain](https://github.com/speechbrain/speechbrain/tree/develop/speechbrain), heavily tied with [HyperPyYAML](https://github.com/speechbrain/HyperPyYAML); released on [PyPI](https://pypi.org/project/speechbrain/)
* [templates](https://github.com/speechbrain/speechbrain/tree/develop/templates)
* [tools](https://github.com/speechbrain/speechbrain/tree/develop/tools) for non-core functionality
* https://huggingface.co/speechbrain/
* hosting several model cards (pretrained models with code snippets)
* Gdrive
* hosting training results; checkpoints; ...

# Conference Tutorials
SpeechBrain has been presented at Interspeech 2021 and 2022 as well as ASRU 2021. When possible, we will provide some ressources here:
- [Interspeech 2022 slides.](https://drive.google.com/drive/folders/1d6GAquxw6rZBI-7JvfUQ_-upeiKstJEo?usp=sharing)
Expand Down
290 changes: 290 additions & 0 deletions docs/coverage.md

Large diffs are not rendered by default.

210 changes: 210 additions & 0 deletions docs/guidance.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,210 @@
# Guiding contributors, reviewers & maintainers through the complexity of SpeechBrain testing.

SpeechBrain is the name of a speech technology toolkit. It is written in Python and uses an extended YAML (HyperPyYAML) for hyperparameters in recipes, as well as some tutorials and scripts. SpeechBrain (the toolkit) is continuously updated and improved by the SpeechBrain community and the SpeechBrain core team, working together on GitHub. New versions of SpeechBrain (the toolkit) are continuously published by the core team on platforms like PyPI.

If we take a step back, SpeechBrain also refers to a wider ecosystem, which has spread to many different platforms: there's documentation on readthedocs, tutorials on Colab, models on HuggingFace, et cetera. Another important part of SpeechBrain are the recipes. The main GitHub repository houses a set of recipes, which has built up over time.

As SpeechBrain (all of it) is improved and changed, ideally the old, existing parts should continue to work well. However, in reality, changes will break old parts.

The purpose of tests is to ensure that things work, or that at least we know what breaks: for example, SpeechBrain (the toolkit) has unittests which test specific bits of code in the core library. But since SpeechBrain (the ecosystem) is quite wide and spread out, there should also be other types of tests which ensure that the different platforms cooperate and the recipes keep working.

Demonstrating that no harm is done by some given change is a big challenge. Ideally, tests will help in integrating (potentially legacy-breaking) changes without losing the existing achievements.

The following graphics illustrate the different complexities at work when it comes to testing in SpeechBrain.

by Andreas Nautsch, Aku Rouhe, 2022

## Functionality provided on multiple platforms, in the SpeechBrain ecosystem.

```
(documentation) (tutorials)
.—————————————. .———————.
| readthedocs | ‚––> | Colab |
\—————————————/ ∕ \———————/
^ ‚––––‘ |
(release) | ∕ v
.——————. .———————————. (landing) .———————————.
| PyPI | –––> | github.io | (page) | templates | (reference)
\——————/ \———————————/ ‚–> \———————————/ (implementation)
| | ‚–––‘ |
v v ∕ v
.———————————–—. .———————————–—. .—————————. .~~~~~~~~~~~~~.
| HyperPyYAML |~~~| speechbrain | ––––––––> | recipes | ––––––––> | HuggingFace |
\————————————–/ \————————————–/ \—————————/ ∕ \~~~~~~~~~~~~~/
(usability) (source/modules) (use cases) ∕ (pretrained models)
| | ∕ |
v v ∕ v
.~~~~~~~~~~~~~. .~~~~~~~~. .———————————.
| PyTorch | ––––––––-> | GDrive | | Inference |
\~~~~~~~~~~~~~/ \~~~~~~~~/ \———————————/
(checkpoints) (results) (code snippets)
```
Each platform/functionality has their own dependencies (which can break) and interfaces (which are specific and can change).

## How is functionality provided?

```
(imported) (used in) (as units in) (integrated by) (to code)
.——————————————. .—————————. .—————————. .—————————. .—————————. | code & yaml
| dependencies | => | helpers | => | classes | => | modules | => | scripts | | style checks
\——————————————/ \—————————/ \—————————/ \—————————/ \—————————/ | (linters)
| | | | |
v v v v v
version updates docstring unittests integration tutorials,
may change their examples as assert tests ensure templates,
interface; modular tests expected working vanilla recipes &
latest versions behaviour experiments snippets
are controlled by need advanced
requirements configs testing strategies

[irregular] [ --- github push workflow actions --- ] [ hybrid periodicity]
```
Python, business as usual:
* doc tests: one or two examples, that the interface does not crash when being used
* unit tests: set of examples to (more exhaustively) test that function does as should
* integration tests: combination of python snippets, targeted yaml hparams, and minimal examples (audio with text & annotation) to demonstrate a use case for a part of module

1. _Contributor, did you provide a new interface?_ <br/>=> doc test
2. _Contributor, did you improve upon inner workings?_ <br/>=> unit test
3. _Contributor, did you offer new ways to the SpeechBrain community?_ <br/>=> integration test

While one cannot control others (dependencies), CI/CD workflows are periodic actions to assert functionality of the known.
Multi-platform checks, for it goes beyond this repo, is on a hybrid (partly irregular periodicity), i.e., before a future SpeechBrain release.

_Naturally, writing style (linters checks) is a part of functionality._

## How is the SpeechBrain community improving quality, continuously?

```
.———————————.
| Closed PR | (but not merged)
‚->\———————————/<-˛
∕ \ (made it! :)
.——————————. .—————————. .———————————.
| Draft PR | –––––––––––> | Open PR | –––––––––––> | Merged PR |
\——————————/ \—————————/ \———————————/
* create initial * ensure all * pre-release
branch to improve workflow checks (later)
* state todo list tests pass * contribution
and fulfill it * collaborate log entry
* inquire feedback on change * part of next
early on requests release tag

| | (more below)
v v

To push formatted code: Review of:
git add ... * changes to core modules
pre-commit * enhanced testing/documentation
git status * contributed tutorial
git add ... * new/edited template/tool
git commit -m ... * added/modified recipe
git push * uploaded pretrained model
* well-formatted py & yaml files
Missed out on one?
pre-commit run --all-files
git status
git add ...
```

To guide the lifecycle of a PR within the SpeechBrain lifecycle—as contributor and as reviewer—can be demanding to being exhausted.
Test automation (e.g., through github and offline workflows) simplify discussions to the points that are of debate, actually.

## The location of a change foreshadows its integrative complexity.

```
BEFORE
------
(python) (yaml)

def func_sig(x, arg0, arg1=None): | my_var: !new:func_sig
# just to demonstrate changes | arg0: 6.28 # tau
if arg1 is None: |
return x + arg0 | my_other: !new:func_sig
else: | arg0: !ref <my_var>
return x + arg1 | arg1: 1/137 # fine structure constant


AFTER - A. Changes to function body &/or interface parameterization via YAML
-----
(python) (yaml)

def func_sig(x, arg0, | my_arg: !new:func_sig
arg1=None,): | arg0: 6.28
if arg1 is None: |
return x / arg0 | my_other: !new:func_sig
else: | arg0: !ref <my_arg>
return x - arg1 | arg1: 0.0073


AFTER - B. Changes to function signature (interface), legacy-preserving
-----
(python) (yaml)

def func_sig(x, arg0, arg1=None, | my_arg: !new:func_sig
arg2=true,): | arg0: 6.28
return next_gen(x, arg0=arg0, |
arg1=arg1, | my_other: !new:func_sig
arg2=arg2,) | arg0: !ref <my_arg>
| arg1: 0.0073
# the new interface being introduced |
def next_gen(x, arg0=6.28, | my_arg_same: !new:next_gen
arg1=1/137, | arg1: None
arg2=true,): |
if !arg2: | my_other_same: !new:next_gen
return x |
if arg1 is None: | my_new_feature: !new:next_gen
return x / arg0 | arg0: 2.718 # e
else: | arg1: 1.618 # what could it be...
return x - arg1 | arg2: false # ;-)


AFTER - C. Changes to function signature (interface), legacy-breaking
-----
(python) (yaml)

def next_gen(x, arg0=6.28, | my_arg: !new:next_gen
arg1=1/137, | arg1: None
arg2=true,): |
if !arg2: | my_other: !new:next_gen
return x |
if arg1 is None: | my_new_feature: !new:next_gen
return x / arg0 | arg0: 2.718
else: | arg1: 1.618
return x - arg1 | arg2: false

```
How would you approach testing each of them?
<br/>Such changes happen not only once, but on a regular basis, throughout all core modules.

Changes can be internal to a function &/or alter the function signature:
* function-internal changes are not of concern to other function (so long they do what they should),
* function signature changes impact the overall—the multi-platform ecosystem.

Legacy-breaking changes will impact the outline of all recipes:
<br/>how will all work after a change—, and after the next major refactoring (after that first one)?

## Branch topology: release <- CI/CD <- ecosystem-spanning refactorings.
```
release | main | business
CI/CD | \--- develop | as usual
ecosystem | \ \<~> testing-refactoring | the tricky
refactoring | \--- unstable <~>/ | bits & pieces
```
The core challenge—to testing SpeechBrain's community-driven development in its multi-platform setting—is tackled through different branches serving each their constructive purpose:
* `main` branch: released on PyPI
* `develop` branch: CI/CD with github workflow; place to merge regular PRs
* `testing-refactoring` branch: [copy of custom interfaces & yaml hparams for pretrained models](https://github.com/speechbrain/speechbrain/tree/testing-refactoring/updates_pretrained_models) hosted on [HuggingFace](https://huggingface.co/speechbrain/)—if changes to the (usually) permanent interface constitution are necessary, we can treat them here (see what happens & improve further)
* `unstable` branch: accumulating legacy-breaking PRs to separate either CI/CD tracks (develop & this one) from one another. When the time of merger comes, the latest `develop` version becomes the final minor release of the passing major version family (e.g., a `0.5.42` before a `0.6.0`). Then, the next lifecycle continues and roots community growth, prepared for new challenges to come.

1. _Contributor, if your change touches upon standing interfaces, then your PR to `develop` or to `unstable` benefits from a companion PR to the `testing-refactoring` branch._<br/>=> Then, reviewers of your main PR is accompanied by provision to also change repos that provide pretrained models.
2. _Contributor, if your idea for change will change function signatures, then your PR strategy needs planning._
1. _Can the change be split into one legacy-preserving (to `develop`) & anoether legacy-breaking PR (to `unstable`)?_<br/> => Then, reviewers of your legacy-preserving PR can help you with facilitating a smooth transition.
2. _Can the legacy-breaking PR be tested for its effectiveness with tools available on the `develop` branch?_ <br/> => Then, reviewers will have their time free to discuss with you on improving your change; provide them tools and assistance to engage with your ideas in a way their mind is open to accept your contribution to the SpeechBrain community.

The other files in this folder provide further guidance on where is what configured, and which tools are there to be used.
Keep in mind, the SpeechBrain community is in-flux, so is a constellation of maintainers and reviewers nothing more but a snapshot.

_Note: github workflows take the definition of a PR, what is specified within its branch. We might update our procedures on the `develop` branch (e.g., to meet dependency updates).
Consequentially, PR and `unstable` branches need to fetch from latest `develop` when testing related definitions are updated._
8 changes: 5 additions & 3 deletions docs/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ or the official `Website <https://speechbrain.github.io>`


License
--------
-------

SpeechBrain is released under the Apache license, version 2.0. The Apache license is a popular BSD-like license.
SpeechBrain can be redistributed for free, even for commercial purposes, although you can not take off the license headers (and under some circumstances you may have to distribute a license document).
Expand All @@ -25,7 +25,7 @@ It is a community project, which means that discussions are engaged community-wi
There is no legal institution associated as an owner of SpeechBrain. Furthermore, and due to the Apache Licence, anyone that would disagree with the way the project is being run can fork it and start a new toolkit.

Referencing SpeechBrain
--------
-----------------------
.. code-block:: txt

@misc{speechbrain,
Expand All @@ -47,9 +47,11 @@ Referencing SpeechBrain
multigpu.md
tutorials.md
contributing.md
guidance.md
coverage.md

API Documentation
--------
-----------------

.. toctree::
:caption: API Documentation:
Expand Down
2 changes: 1 addition & 1 deletion pytest.ini
Original file line number Diff line number Diff line change
Expand Up @@ -6,4 +6,4 @@ python_files =
check_*.py
example_*.py

norecursedirs = results
norecursedirs = results tmp utils
3 changes: 2 additions & 1 deletion recipes/AISHELL-1/ASR/CTC/hparams/train_with_wav2vec.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ valid_data: !ref <output_folder>/dev.csv
test_data: !ref <output_folder>/test.csv

wav2vec2_hub: TencentGameMate/chinese-wav2vec2-large
wav2vec2_folder: !ref <save_folder>/wav2vec2_checkpoint

# Training parameters
number_of_epochs: 80
Expand Down Expand Up @@ -130,7 +131,7 @@ wav2vec2: !new:speechbrain.lobes.models.huggingface_wav2vec.HuggingFaceWav2Vec2
source: !ref <wav2vec2_hub>
output_norm: True
freeze: !ref <freeze_wav2vec>
save_path: !ref <save_folder>/wav2vec2_checkpoint
save_path: !ref <wav2vec2_folder>

ctc_lin: !new:speechbrain.nnet.linear.Linear
input_size: !ref <dnn_neurons>
Expand Down
Loading