AgentPantheon
Relari (YC W24) logo

Relari (YC W24)AI araçları için test, değerlendirmesi ve sentetik veri oluşturma platformu.

4.3 (6)
Daniel Nikulshynİnceleyen Daniel Nikulshyn·Güncellendi Temmuz 2026

Genel Bakış

Relari, YC W24, geliştirici platformu, AI ajanlarının güvenilirliğini geliştirmek için sistematik testleme ve değerlendirme üzerinde odaklanmaktadır. Takımları oluşturulmuş datasetler oluşturmayı, otomatik değerlendirmeleri işleyip ve gerçekçi senaryolarda ajan performansını ölçümleyerek üretim ortamına göndermeden önce göndermektedir. Y Combinator'ın (W24) desteklediği Relari, geleneksel kalite güvencesinin yetersiz olduğu karmaşık LLM uygulamaları ve sıralı agresler inşa eden mühendis团lar için tasarlanmıştır. Ardışıl QA'nın sınırlarının ötesinde bulunan araçları, deterministik olmayan AI sistemlerine yazılım mühendisliği kuralları ile beraber - birim testleri, geri kazanım kontrolleri ve ölçülebilir istatistikler - tanıtmayı amaçlar. Platform, özel değerlendiriciler, senaryo simülasyonu ve sürekli izleme desteği sunuyor, böylece hem önceden yayınlanmadan önce doğrulama hem de üretim.agentlarının kalite güvencesi sağlama gibi kullanım durumlarının destekleniyor.

Temel özellikler

  • Sentetik veri kümesi oluşturma
  • Otomatik agent değerlendirme.pipeline'leri
  • Senaryo ve konuşma simülasyonu
  • Özenli değerlendirme ölçütleri
  • Geliştirilmiş LLM uygulamaları için gerileme testleri
  • Sürat performansı ve raporlama için benchmarking

Fiyatlar

Model
Free
Puan
4.3 / 5 (6)

Kullanım senaryoları

AI araçlarının test edilmesi

Güvenilir ve test edilmüş AI araçları değerlendirmesi için.

Artılar ve eksiler

Artılar

  • Çok adımlı AI araçları için özel olarak tasarlandı.
  • Düzenli sentetik test verisi üretmek için ölçekli.
  • Özel ölçütler ve değerlendiriciler desteklenmektedir.
  • Y Combinator tarafından desteklenmekte ve aktif geliştirme sürdürülmektedir.

Eksiler

  • Primarily ana olarak teknik takımlar için yönlendirilmiştir, geliştiriciler olmayan için değil.
  • Yenilikçi bir platform olduğu için özellikleri bir evrim geçiriyor.
  • Varolan stacklar için uyum için entegrasyon çalışmaları gerekebilir.

İncelemeler

4.3

6 puandan ortalama.

5
2
4
4
3
0
2
0
1
0

İnceleme bırakmak için giriş yap.

F

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

R

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

D

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

C

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Y

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

L

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Sorular

Henüz soru yok — ilk soruyu sen sor.

Soru sor

Observability alternatifleri