arrow_backブログに戻る
AI Companion Memory Test: A 7-Day Benchmark for Recall and Continuity

AI Companion Memory Test: A 7-Day Benchmark for Recall and Continuity

Run a repeatable 7-day benchmark to test AI companion memory recall, correction handling, and event continuity across any app.

Elyvie 編集チーム
公開日 2026年8月1日·更新日 2026年8月1日·7 分で読めます
ai companion memoryai memory testchatbot memory benchmarkai companion comparisonconversation continuity

Most “best AI companion memory” lists compare feature pages. That tells you whether an app has a memory button, but not whether a companion can recall the right detail at the right moment, update it when life changes, and avoid mixing it with someone else’s story.

This guide gives you a repeatable seven-day AI companion memory test. It does not declare a universal winner or pretend that published product claims are independent test results. Instead, it helps you test any companion using the same facts, timing, corrections, and scoring criteria.

What good AI companion memory actually means

A companion does not have strong memory simply because it can repeat your name. Useful memory has several parts:

  1. Fact recall: It remembers a stable preference or personal detail.
  2. Event continuity: It follows up on something that was going to happen.
  3. Correction handling: It replaces outdated information instead of repeating both versions.
  4. Relationship continuity: It remembers shared moments without turning every conversation into a database lookup.
  5. Boundaries and control: You can understand, correct, or delete what the app remembers.

Short-term context can make an app look excellent during one long conversation. The harder test is whether the right information returns after time has passed or a new chat has started.

Before the test: keep the setup fair

Use a new companion or a clean conversation where possible. Pick five harmless fictional test facts rather than sensitive real information.

For example:

  • Your favorite drink is roasted oolong tea.
  • Your dog is named Pepper and dislikes thunderstorms.
  • Your sister Maya is visiting on Sunday.
  • You have a presentation on Friday morning.
  • You are learning pottery and struggle with centering the clay.

These are testing facts, not real users or testimonials. Use the same wording and timing in every app you compare. Do not repeatedly remind the companion of the answers, because that tests your prompting rather than its memory.

Also record the plan, app version, and subscription tier. Memory features can differ between free and paid tiers and can change over time.

Day 1: establish the facts naturally

Do not paste the five facts as a checklist. Work them into a normal conversation across several messages.

You might mention making oolong tea while talking about your day, then explain that Pepper hides during storms. Later, mention Maya’s Sunday visit and the Friday presentation. End by talking about the pottery problem.

Afterward, check whether the app exposes a memory screen. Note which facts were saved automatically, which require manual pinning, and whether you can edit them.

A visible memory entry is useful evidence, but it is not the same as successful recall. The next six days test whether those memories affect conversation.

Day 2: test unprompted factual recall

Start a fresh conversation and ask open questions that do not contain the answer:

  • “What drink would you make for me?”
  • “What tends to upset my dog?”
  • “What am I trying to learn lately?”

Give one point for each correct answer. Give half a point if the companion needs a broad hint, and zero if you supply the answer inside the question.

Watch for confident inventions. “I don’t remember” is a better failure mode than confidently giving your dog the wrong name.

Day 4: test corrections and conflicts

Tell the companion that one fact has changed:

I still like oolong, but lately coffee has become my morning drink.

Later that day, ask what you usually drink in the morning. A good system should preserve the distinction: oolong remains a preference, while coffee is the newer routine.

Score two things separately:

  • Did it learn the new fact?
  • Did it avoid erasing or mangling the older context?

This matters because real lives change. A memory system that only accumulates facts eventually creates contradictions.

Day 6: test shared-event continuity

Ask the companion what is happening later in the week. Do not mention Maya, Sunday, or the presentation.

Strong event continuity should surface the upcoming presentation or visit when relevant. It should not force the detail into an unrelated conversation merely to prove that it remembers.

Give one point for accurate recall and one point for appropriate timing.

Day 7: test follow-up after the event

Say that Friday has passed, then begin a normal conversation. Does the companion ask how the presentation went? If it does, provide a result such as:

It went well, but the questions at the end were difficult.

Later, ask what happened with the presentation. This tests whether the system can connect the planned event with its outcome rather than storing two disconnected facts.

A simple 10-point memory score

Use the same scorecard for every app:

  • Two points for Day 2 factual recall
  • Two points for correction handling
  • Two points for event continuity
  • Two points for following up and connecting an outcome
  • One point for admitting uncertainty instead of inventing
  • One point for memory controls: view, edit, or delete

Keep a separate note for conversational quality. A companion can retrieve facts accurately and still use them awkwardly. The best AI companion memory for you is the one that combines recall with natural timing and gives you acceptable control over your data.

How current products describe their memory systems

Official documentation shows that products take different approaches. These descriptions are product documentation, not the result of this benchmark.

Character.AI describes Story Memory, automatically captured Facts, and pinned information that users can manage. Its May 2026 update says Facts can cover the user persona, the Character, and side characters, with options to edit or remove them. See Character.AI’s memory update.

Replika describes layered memory, including a visible Memory tab and deeper personalization based on conversation patterns. Users can also add or remove certain memories manually. See Replika’s memory explanation.

Kindroid documents persistent, cascaded, and retrievable memory spread across several systems. Its documentation also distinguishes paid-tier context from long-term retrieval and notes that retrieved memories may not always be recalled. See Kindroid’s memory guide.

Nomi describes short-, medium-, and long-term memory alongside an Identity Core that updates around important facts, preferences, feedback, and shared experiences. See Nomi’s Identity Core explanation.

The useful question is not which product uses the most impressive label. It is which design produces reliable, appropriate recall in your actual conversations.

Questions to ask before trusting a memory feature

Before sharing anything personal, check:

  • Can you see what was saved?
  • Can you correct or delete individual memories?
  • Does deleting a chat also delete extracted memories?
  • Are memories isolated between different companions?
  • Does a new chat carry memories forward automatically?
  • Are memory features different on free and paid plans?
  • Is there a temporary or no-memory conversation mode?

Read the current privacy policy rather than assuming that “memory” means the same storage model everywhere.

Testing Elyvie with the same benchmark

Elyvie uses recent conversation context plus structured long-term facts to support cross-session continuity. Memories are separated by companion so information shared with one character is not intended to become another character’s knowledge. The product explanation is available in How Elyvie Works.

That architecture should be tested with the same benchmark rather than exempted because this article appears on Elyvie’s site. Start a conversation, use the five fictional facts, and judge whether the follow-up feels accurate and natural.

The answer: test recall, not feature names

Which AI companion remembers conversations best? There is no honest universal answer without a defined scenario, a time interval, and a consistent test.

Run the same seven-day benchmark. Keep the facts harmless. Score corrections and uncertainty, not only successful callbacks. The result will tell you much more than a marketing page that simply promises “infinite memory.”

Sources

How this article was produced

This article compares current official product documentation and proposes a fictional, repeatable benchmark. Elyvie did not conduct a controlled head-to-head product trial for this article, and no fictional test facts represent real users.

Elyvie をもっと見る

関連記事

当サイトでは、ユーザー体験向上のためにCookieを使用しています。