v0.2.4

yifanmai released this 20 Sep 18:48

· 825 commits to main since this release

Models

Added Meta LLaMA, Meta Llama 2, EleutherAI Pythia, Together RedPajama on Together (#1821)
Removed the unofficial chat-gpt client in favor of the official API (#1809)
Added support for models for the NeurIPS Efficiency Challenge (#1693)

Frontend

Added support for rendering train-test overlap stats in the frontend (#1747)
Fixed a bug where stats with NaN values would cause the frontend to fail to render tables (#1784)

Framework

Moved many dependencies, especially those only used by a single model provider or a small number of runs, to optional extra dependencies (#1798, #1844)
Widened some dependencies (e.g. PyTorch) to reduce dependency conflicts with other packages (#1759)
Added MaxEvalInstancesRunExpander to allow overriding the number of eval instances at the run level (#1837)
Updated human critique evaluation on Amazon Mechanical Turk to support emoji and other special characters (#1773)
Fixed a bug where in-context learning examples with multiple correct references were adapted to prompts where all the correct references are concatenated together as the output, which was not intended for some scenarios (e.g. narrative_qa, natural_qa, quac and wikifact) (#1785)
Fixed a bug where ObjectSpec is not hashable if any arg is a list (#1771)

Evaluations

Added evaluation results for Meta LLaMA, Meta Llama 2, EleutherAI Pythia, Together RedPajama on Together
Corrected evaluation results for AI21 Jurassic-2 and Writer Palmyra for the scenarios narrative_qa, natural_qa, quac and wikifact, as they were affected by the bug fixed by #1785

Contributors

Thank you to the following contributors for your contributions to this HELM release!

Contributors

yifanmai, percyliang, and 9 other contributors

Assets 2