A pre-trained model with multi-exit transformer architecture.

Last update: Dec 14, 2022

Related tags

Deep Learning ElasticBERT

Overview

ElasticBERT

This repository contains finetuning code and checkpoints for ElasticBERT.

Towards Efficient NLP: A Standard Evaluation and A Strong Baseline

Xiangyang Liu, Tianxiang Sun, Junliang He, Lingling Wu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, Xipeng Qiu

Requirements

We recommend using Anaconda for setting up the environment of experiments:

conda create -n elasticbert python=3.8.8
conda activate elasticbert
conda install pytorch==1.8.1 cudatoolkit=11.1 -c pytorch -c conda-forge
pip install -r requirements.txt

Pre-trained Models

We provide the pre-trained weights of ElasticBERT-BASE and ElasticBERT-LARGE, which can be directly used in Huggingface-Transformers.

ElasticBERT-BASE: 12 layers, 12 Heads and 768 Hidden Size.
ElasticBERT-LARGE: 24 layers, 16 Heads and 1024 Hidden Size.

The pre-trained weights can be downloaded here.

Model	`MODEL_NAME`
`ElasticBERT-BASE`	fnlp/elasticbert-base
`ElasticBERT-LARGE`	fnlp/elasticbert-large

Downstream task datasets

The GLUE task datasets can be downloaded from the GLUE leaderboard

The ELUE task datasets can be downloaded from the ELUE leaderboard

Finetuning in static usage

We provide the finetuning code for both GLUE tasks and ELUE tasks in static usage on ElasticBERT.

For GLUE:

cd finetune-static
bash finetune_glue.sh

For ELUE:

cd finetune-static
bash finetune_elue.sh

Finetuning in dynamic usage

We provide finetuning code to apply two kind of early exiting methods on ElasticBERT.

For early exit using entropy criterion:

cd finetune-dynamic
bash finetune_elue_entropy.sh

For early exit using patience criterion:

cd finetune-dynamic
bash finetune_elue_patience.sh

Please see our paper for more details!

Contact

If you have any problems, raise an issue or contact Xiangyang Liu

Citation

If you find this repo helpful, we'd appreciate it a lot if you can cite the corresponding paper:

@article{liu2021elasticbert,
  author    = {Xiangyang Liu and
               Tianxiang Sun and
               Junliang He and
               Lingling Wu and
               Xinyu Zhang and
               Hao Jiang and
               Zhao Cao and
               Xuanjing Huang and
               Xipeng Qiu},
  title     = {Towards Efficient {NLP:} {A} Standard Evaluation and {A} Strong Baseline},
  journal   = {CoRR},
  volume    = {abs/2110.07038},
  year      = {2021},
  url       = {https://arxiv.org/abs/2110.07038},
  eprinttype = {arXiv},
  eprint    = {2110.07038},
  timestamp = {Fri, 22 Oct 2021 13:33:09 +0200},
  biburl    = {https://dblp.org/rec/journals/corr/abs-2110-07038.bib},
  bibsource = {dblp computer science bibliography, https://dblp.org}
}

A pre-trained model with multi-exit transformer architecture.

Related tags

Overview

ElasticBERT

Requirements

Pre-trained Models

Downstream task datasets

Finetuning in static usage

Finetuning in dynamic usage

Contact

Citation

Owner

fastNLP

Automatic number plate recognition using tech: Yolo, OCR, Scene text detection, scene text recognation, flask, torch

CausaLM: Causal Model Explanation Through Counterfactual Language Models

Code for EmBERT, a transformer model for embodied, language-guided visual task completion.

Python Implementation of the CoronaWarnApp (CWA) Event Registration

Official page of Struct-MDC (RA-L'22 with IROS'22 option); Depth completion from Visual-SLAM using point & line features

This is the pytorch implementation for the paper: Learning Accurate Performance Predictors for Ultrafast Automated Model Compression, which is in submission to TPAMI

A set of examples around hub for creating and processing datasets

Code for reproducing our analysis in the paper titled: Image Cropping on Twitter: Fairness Metrics, their Limitations, and the Importance of Representation, Design, and Agency

Official PyTorch Implementation of SSMix (Findings of ACL 2021)

SimulLR - PyTorch Implementation of SimulLR

HGCN: Harmonic Gated Compensation Network For Speech Enhancement

AirLoop: Lifelong Loop Closure Detection

A DCGAN to generate anime faces using custom mined dataset

C3d-pytorch - Pytorch porting of C3D network, with Sports1M weights

Code for ACL 2019 Paper: "COMET: Commonsense Transformers for Automatic Knowledge Graph Construction"

my graduation project is about live human face augmentation by projection mapping by using CNN

Joint Learning of 3D Shape Retrieval and Deformation, CVPR 2021

ByteTrack(Multi-Object Tracking by Associating Every Detection Box)のPythonでのONNX推論サンプル

A coin flip game in which you can put the amount of money below or equal to 1000 and then choose heads or tail

Semantic Segmentation for Aerial Imagery using Convolutional Neural Network

A pre-trained model with multi-exit transformer architecture.

Related tags

Overview

ElasticBERT

Requirements

Pre-trained Models

Downstream task datasets

Finetuning in static usage

Finetuning in dynamic usage

Contact

Citation

Owner

fastNLP

Automatic number plate recognition using tech: Yolo, OCR, Scene text detection, scene text recognation, flask, torch

CausaLM: Causal Model Explanation Through Counterfactual Language Models

Code for EmBERT, a transformer model for embodied, language-guided visual task completion.

Python Implementation of the CoronaWarnApp (CWA) Event Registration

Official page of Struct-MDC (RA-L'22 with IROS'22 option); Depth completion from Visual-SLAM using point & line features

This is the pytorch implementation for the paper: *Learning Accurate Performance Predictors for Ultrafast Automated Model Compression*, which is in submission to TPAMI

A set of examples around hub for creating and processing datasets

Code for reproducing our analysis in the paper titled: Image Cropping on Twitter: Fairness Metrics, their Limitations, and the Importance of Representation, Design, and Agency

Official PyTorch Implementation of SSMix (Findings of ACL 2021)

SimulLR - PyTorch Implementation of SimulLR

HGCN: Harmonic Gated Compensation Network For Speech Enhancement

AirLoop: Lifelong Loop Closure Detection

A DCGAN to generate anime faces using custom mined dataset

C3d-pytorch - Pytorch porting of C3D network, with Sports1M weights

Code for ACL 2019 Paper: "COMET: Commonsense Transformers for Automatic Knowledge Graph Construction"

my graduation project is about live human face augmentation by projection mapping by using CNN

Joint Learning of 3D Shape Retrieval and Deformation, CVPR 2021

ByteTrack(Multi-Object Tracking by Associating Every Detection Box)のPythonでのONNX推論サンプル

A coin flip game in which you can put the amount of money below or equal to 1000 and then choose heads or tail

Semantic Segmentation for Aerial Imagery using Convolutional Neural Network

This is the pytorch implementation for the paper: Learning Accurate Performance Predictors for Ultrafast Automated Model Compression, which is in submission to TPAMI