AgentPantheon
Relari (YC W24) logo

Relari (YC W24)תפיסה, הערכה וייצור נתונים סינתטיים לאגנטי AI

4.3 (6)
Daniel Nikulshynנבדק על ידי Daniel Nikulshyn·עודכן יולי 2026

סקירה

Relari היא פלטפורמה למפתחים המתמקדת בשיפור מהימנותם של סוכני AI באמצעות בדיקה והערכה שיטתיים. היא עוזרת לצוותים ליצור מערכי נתונים סינתטיים, לבצע הערכות אוטומטיות ולמדוד את ביצועי הסוכנים במגוון תרחישים מציאותיים לפני הפצה לייצור. בגיבוי של Y Combinator (W24), רלרי מתמקדת בצוותים הנדסיים הבונים יישומי LLM מורכבים וסוכנים רב-שלביים שבהם בקרת האיכות המסורתית אינה מספיקה. הכלים שלה שואפים להביא קפדנות של הנדסת תוכנה - בדיקות יחידה, בדיקות רגרסיה ומדדים ניתנים למדידה - למערכות AI לא דטרמיניסטיות. הפלטפורמה תומכת במעריכים מותאמים אישית, הדמיית תרחישים ומעקב רציף, מה שהופך אותה לשימושית הן לאימות טרום-השקה והן להבטחת איכות שוטפת של סוכנים בייצור.

תכונות עיקריות

  • ייצור את אסמבלים סינתטיים
  • הערכה בבירוקרטיה של האגנטים
  • אימולירת סצנריות ושיחות
  • מדדי ביקורת-הערכה ניתנת להגדרה
  • תביר-השוואה-הדדיה לרכזי LLM
  • איפוסד מדדים ודיווחי ביצוע
  • regression testing for LLM

תמחור

מודל
Free
קטגוריה
Observability
דירוג
4.3 / 5 (6)

מקרי שימוש

TESTING AGNENTI AI

אומד-איכות-ראיה-אוטומטית של AGNENTI AI

יתרונות וחסרונות

יתרונות

  • מעוצב באופן נפשי ה-LLM
  • מוצא נתונים-בד-סינתטיים בהיקף
  • שמור והתאמתי את סטטיסטי-מדדים ואבלים
  • בעלת פעילות אקטיבית ופיתוח מטות Y Combinator
  • may
  • primarily aimed at technical teams

חסרונות

  • מיועד בעיקר לצוותים טכניים, לא למפתחים שאינם
  • פלטפורמה חדשה יחסית עם קבוצת תכונות המתפתחת
  • ייתכן שיהיה צורך בעבודה לשילוב כדי להתאים לערימות ק существующие

ביקורות

4.3

ממוצע מ-6 דירוגים.

5
2
4
4
3
0
2
0
1
0

התחבר כדי להשאיר ביקורת.

F

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

R

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

D

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

C

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Y

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

L

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

שאלות ותשובות

עדיין אין שאלות — היה הראשון לשאול.

שאל שאלה

חלופות לObservability