AgentPantheon
Relari (YC W24) logo

Relari (YC W24)Platforma na testovanie, hodnotenie a generovanie syntetických dát pre AI agentov.

4.3 (6)
Daniel NikulshynRecenzované Daniel Nikulshyn·Aktualizované júl 2026

Prehľad

Relari je platforma pre vývojárov zameraná na zlepšenie spoľahlivosti AI agentov prostredníctvom systematického testovania a hodnotenia. Pomáha tímom vytvárať syntetické súbory údajov, spúšťať automatizované hodnotenia a porovnávať výkonnosť agentov v realistických scenároch pred nasadením do produkcie. Spoločnosť Relari, podporovaná spoločnosťou Y Combinator (W24), sa zameriava na tímy vývojárov, ktoré vytvárajú zložité aplikácie LLM a viacstupňové agenty, kde tradičné testovanie kvality nestačí. Ich nástroje majú za cieľ priniesť prísnosť softvérového inžinierstva - jednotkové testy, regresné kontroly a merateľné metriky - do nedeterministických systémov AI. Platforma podporuje vlastné evaluátory, simuláciu scenárov a priebežné monitorovanie, čo ju robí užitočnou pre overenie pred spustením aj pre priebežné zabezpečovanie kvality produkčných agentov.

Kľúčové funkcie

  • Generovanie syntetických dátových sád
  • Automatizované hodnotiace pipeline pre agentov
  • Simulácia scenárov a konverzácií
  • Prispôsobiteľné hodnotiace metriky
  • Regresné testovanie pre LLM aplikácie
  • Benchmarkovanie výkonnosti a reportovanie

Cenník

Model
Free
Kategória
Observability
Hodnotenie
4.3 / 5 (6)

Prípady použitia

Testovanie AI agentov

Spoľahlivé a testovateľné hodnotenie pre AI agentov.

Klady a zápory

Klady

  • Navrhnuté špeciálne na hodnotenie viackrokových AI agentov
  • Generuje syntetické testovacie dáta vo veľkom meradle
  • Podporuje vlastné metriky a evaluátory
  • Podporovaná Y Combinator s aktívnym vývojom

Zápory

  • Primárne určené pre technické tímy, nie pre ne‑vývojárov
  • Novšia platforma s neustále sa vyvíjajúcim súborom funkcií
  • Môže vyžadovať integračnú prácu na prispôsobenie existujúcim stackom

Recenzie

4.3

Priemer z 6 hodnotení.

5
2
4
4
3
0
2
0
1
0

Prihlás sa, aby si napísal recenziu.

F

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

R

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

D

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

C

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Y

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

L

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Otázky

Žiadne otázky — polož prvú.

Polož otázku

Alternatívy k Observability