> For the complete documentation index, see [llms.txt](https://gitbook.bergelsonlab.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://gitbook.bergelsonlab.com/archive/old-osf-pages/scripts-and-the-like/clanstats.md).

# clanstats

## clanstats

github repo [here](https://github.com/SeedlingsBabylab/clanstats)

### usage

#### [clanstats.py](http://clanstats.py/)

This script will output the long form speaker/classifier data when given a single .cex clan file. It takes 3 arguments. The .cex file (first argument) can also be replaced by the .csv output produced by parse\_clan. If the input is a parse\_clan csv file, the window size will be 0.

```
$: python clanstats.py  /path/to/clanfile.cex   /path/to/output   window_size
```

#### cs\_folder.py

This script will run [clanstats.py](http://clanstats.py/) on every clan file (or parse\_clan .csv) in a directory passed as argument.

```
$: python cs_folder.py  /path/to/clanfiles/   /output/path/   window_size
```

If there are parse\_clan csv files in the folder being batch processed, the window\_size will automaticaly be set as 0. This means the clan files will be processed with the window, while the csv's will not. If you need window size consistency across all the outputs, use a window\_size of 0, or leave the csv files out of the folder being batch processed.

This script also produces a single .csv file containing the aggregate data from all the files it just processed (named aggregate\_long.csv). It'll combine all the .csv's in the output directory (that was originally passed as an argument), so if you have .csv files in there that weren't a result of what cs\_folder.py just did, it'll combine those as a part of the aggregate\_long.csv output as well. Make sure there's only .csv files in the output directory or else the script will throw an error when trying to concatenate those files.

**csv file input**

The csv input that clanstats accepts should be in the form produced by the parse\_clan2 script, directions found [here](https://osf.io/4rcjd/wiki/ParseClan2/). The github repo is [here](https://github.com/SeedlingsBabylab/parse_clan2).

There's also a helper script in the parse\_clan2 folder called "batch\_parse\_clan2.py". This is helpful if you have a folder filled with .cha files and you want to run parse\_clan2 over all of them.

**batch\_parse\_clan2.py usage**

```
$: python batch_parse_clan2.py /folder/with/cha/files  /output/folder
```
