Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LongMemEvalS

Benchmarks

Task NameDataset NameSOTA ResultTrend
Long-context Memory Retrieval and QALongMemEvalS
Accuracy61.67
12
Relevance RankingLongMemEvalS
Macro-AUC0.752
8
Question AnsweringLongMemEvalS
Accuracy33.8
5
Showing 3 of 3 rows