AgentPantheon
Relari (YC W24) logo

Relari (YC W24)Platforma do testowania, oceny i generowania danych syntetycznych dla agentów AI.

4.3 (6)
Daniel NikulshynZrecenzowane przez Daniel Nikulshyn·Zaktualizowano lipiec 2026

Przegląd

Relari to platforma dla programistów skoncentrowana na zwiększaniu niezawodności agentów AI poprzez systematyczne testowanie i ocenę. Pomaga zespołom generować zestawy danych syntetycznych, uruchamiać automatyczne oceny i porównywać wydajność agentów w realistycznych scenariuszach przed wdrożeniem do produkcji. Wspierany przez Y Combinator (W24), Relari skierowany jest do zespołów inżynieryjnych budujących złożone aplikacje LLM i wieloetapowe agenty, gdzie tradycyjna kontrola jakości zawodna. Narzędzia mają na celu wprowadzenie rygoru inżynierii oprogramowania — testów jednostkowych, kontroli regresji i mierzalnych metryk — do systemów AI o niedeterministycznym zachowaniu. Platforma obsługuje niestandardowe ewaluatory, symulację scenariuszy i ciągłe monitorowanie, co czyni ją przydatną zarówno do weryfikacji przed uruchomieniem, jak i stałej kontroli jakości agentów produkcyjnych.

Kluczowe funkcje

  • Generowanie zestawów danych syntetycznych
  • Automatyczne potoki oceny agentów
  • Symulacja scenariuszy i konwersacji
  • Dostosowywalne metryki oceny
  • Testy regresji dla aplikacji LLM
  • Benchmarking wydajności i raportowanie

Cennik

Model
Free
Kategoria
Observability
Ocena
4.3 / 5 (6)

Zastosowania

Testowanie agentów AI

Niezawodna i testowalna ocena agentów AI

Plusy i minusy

Plusy

  • Specjalnie zaprojektowane do oceny wieloetapowych agentów AI
  • Generuje dane testowe syntetyczne w dużej skali
  • Obsługuje niestandardowe metryki i ewaluatory
  • Wspierane przez Y Combinator z aktywnym rozwojem

Minusy

  • Przede wszystkim skierowane do zespołów technicznych, nie do osób nietechnicznych
  • Nowsza platforma z rozwijającym się zestawem funkcji
  • Może wymagać pracy integracyjnej, aby dopasować do istniejących stosów

Recenzje

4.3

Średnia z 6 ocen.

5
2
4
4
3
0
2
0
1
0

Zaloguj się, aby zostawić recenzję.

F

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

R

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

D

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

C

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Y

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

L

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Pytania i odpowiedzi

Brak pytań — zadaj pierwsze.

Zadaj pytanie

Alternatywy dla Observability