AgentPantheon
Relari (YC W24) logo

Relari (YC W24)Piattaforma per il testing, l'evaluation e la generazione di dati sintetici per gli agenti di intelligenza artificiale.

4.3 (6)
Daniel NikulshynRecensito da Daniel Nikulshyn·Aggiornato luglio 2026

Panoramica

Relari è una piattaforma per sviluppatori focalizzata sul miglioramento dell'affidabilità degli agenti AI attraverso test e valutazioni sistematici. Aiuta i team a generare dataset sintetici, eseguire valutazioni automatizzate e confrontare le prestazioni degli agenti in scenari realistici prima di rilasciarli in produzione. Supportato da Y Combinator (W24), Relari si rivolge a team di ingegneria che sviluppano applicazioni LLM complesse e agenti multi-step in cui il QA tradizionale è insufficiente. I suoi strumenti mirano a portare rigore ingegneristico del software - test unitari, controlli di regressione e metriche misurabili - a sistemi AI non deterministici. La piattaforma supporta valutatori personalizzati, simulazione di scenari e monitoraggio continuo, rendendola utile sia per la convalida pre-lancio che per la garanzia di qualità continua degli agenti di produzione.

Funzionalità chiave

  • Generazione del set di dati sintetico
  • Flussi di valutazione auto-matizzati degli agenti
  • Simulazione dello scenario e della conversazione
  • Metriche di valutazione personalizzabili
  • Test di regressione per gli applicativi LLM
  • Benchmarking e relazione delle prestazioni

Prezzi

Modello
Free
Categoria
Observability
Valutazione
4.3 / 5 (6)

Casi d’uso

Testing degli aggenti di intelligenza artificiale

Valutazione sicura e gestibile degli agenti di intelligenza artificiale.

Pro & contro

Pro

  • Progettato per evausi multi-passaggio degli agenti di intelligenza artificiale
  • Genera dei dati di test sintetici ad scala
  • Supporta metriche e valutatori personalizzabili
  • Sostenuo da parte dell'azienda Y Combinator con sviluppo attivo

Contro

  • Principalmente destinato a team tecnici, non a non sviluppatori
  • Piattaforma più nuova con un insieme di caratteristiche in evoluzione
  • Potrebbe richiedere lavoro di integrazione per adattare la pila esistente

Recensioni

4.3

Media su 6 valutazioni.

5
2
4
4
3
0
2
0
1
0

Accedi per lasciare una recensione.

F

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

R

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

D

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

C

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Y

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

L

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Domande e risposte

Ancora nessuna domanda — sii il primo a chiedere.

Fai una domanda

Alternative a Observability