evals

by openai · Developer Tools

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

New0 ratings19,395 starsActive
Developer ToolsDeveloper ToolPythonHTML
GitHub⚑ Report
evals — image 1

About this project

OpenAI Evals You can now configure and run Evals directly in the OpenAI Dashboard. Get started → Evals provide a framework for evaluating large language models (LLMs) or systems built using LLMs. We offer an existing registry of evals to test different dimensions of OpenAI models and the ability to write your own custom evals for use cases you care about. You can also use your data to build private evals which represent the common LLMs patterns in your workflow without exposing any of that data publicly. If you are building with LLMs, creating high quality evals is one of the most impactful things you can do. Without evals, it can be very difficult and time intensive to understand how…

Technologies

JavaScriptShellPythonHTMLJupyter Notebook

Project health

Low recent activity
Last update4 months ago
Contributors435
Latest release
Open issues & PRs333
LicenseOther
On GitHubsince 2023

GitHub

19,395
stars
3,077
forks
435
contributors
333
open issues & PRs
Python
language
4 months ago
last commit
View on GitHub ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review evals.

Built by

openai
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain openai/evals? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like