
Retentioneering Tutorial for User Behavior Analytics
Understanding how users interact with your website or product is crucial for optimizing conversions and driving business growth. The Retentioneering library empowers analysts and product teams to map, visualize, and analyze user behavior flows in detail, uncovering actionable insights from event data.
What Is Retentioneering and Why Use It?
Retentioneering is an open-source Python library designed for user behavior analytics.
Unlike traditional analytics tools that focus on isolated metrics, Retentioneering lets you explore the entire journey users take through your product by representing actions as event streams.
This approach allows for advanced analyses including funnel conversion, path analysis, and retention cohort studies.
Retentioneering is especially useful for:
- Visualizing user journeys and identifying friction points
- Optimizing conversion funnels
- Comparing behaviors of different user segments
- Diagnosing drop-off and improving retention
Core Concepts for User Behavior Analytics
Before diving into code, let’s review a few core concepts you’ll need:
- Events: Actions users take (e.g.,
product_view,add_to_cart), each logged with a timestamp and user identifier. - Event Streams: Sequences of events for each user, ordered by time.
- Funnels: Prescribed paths (such as
product_view → add_to_cart → purchase) where we measure conversion/drop-off at each step. - Flow Graphs: Visual representations of all user transitions between events, highlighting common paths and loops.
- Retention: The rate at which users return or progress through the product after initial activity, often analyzed via cohort analysis.
Setting Up: Installation and Imports
Let's ensure the necessary packages are installed. Retentioneering and its dependencies can be installed via pip:
pip install retentioneering pandas matplotlib seaborn
Now, import the required libraries in your Python script or notebook:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import retentioneering as re
Example E-Commerce User Journey Dataset
We'll use a synthetic dataset that simulates user interactions on an e-commerce platform. Each record represents a user action such as viewing a product, adding to cart, checking out, or completing a purchase.
Sample Event Data
from datetime import datetime, timedelta
# Create a synthetic event dataset
data = [
# user_id, event_name, timestamp
[1, "product_view", datetime(2024, 6, 1, 12, 0, 0)],
[1, "add_to_cart", datetime(2024, 6, 1, 12, 2, 0)],
[1, "checkout", datetime(2024, 6, 1, 12, 5, 0)],
[1, "purchase", datetime(2024, 6, 1, 12, 7, 0)],
[2, "product_view", datetime(2024, 6, 1, 12, 1, 0)],
[2, "product_view", datetime(2024, 6, 1, 12, 4, 0)],
[2, "add_to_cart", datetime(2024, 6, 1, 12, 5, 0)],
[2, "checkout", datetime(2024, 6, 1, 12, 7, 0)],
[3, "product_view", datetime(2024, 6, 1, 12, 3, 0)],
[3, "product_view", datetime(2024, 6, 1, 12, 6, 0)],
[3, "add_to_cart", datetime(2024, 6, 1, 12, 8, 0)],
[4, "product_view", datetime(2024, 6, 1, 12, 4, 0)],
[4, "add_to_cart", datetime(2024, 6, 1, 12, 6, 0)],
[4, "checkout", datetime(2024, 6, 1, 12, 8, 0)],
[4, "purchase", datetime(2024, 6, 1, 12, 12, 0)],
[5, "product_view", datetime(2024, 6, 1, 12, 5, 0)],
[5, "product_view", datetime(2024, 6, 1, 12, 6, 0)],
[5, "add_to_cart", datetime(2024, 6, 1, 12, 7, 0)],
]
df = pd.DataFrame(data, columns=["user_id", "event", "timestamp"])
df.sort_values(["user_id", "timestamp"], inplace=True)
df.reset_index(drop=True, inplace=True)
df.head(10)
| user_id | event | timestamp |
|---|---|---|
| 1 | product_view | 2024-06-01 12:00:00 |
| 1 | add_to_cart | 2024-06-01 12:02:00 |
| 1 | checkout | 2024-06-01 12:05:00 |
| 1 | purchase | 2024-06-01 12:07:00 |
| 2 | product_view | 2024-06-01 12:01:00 |
| 2 | product_view | 2024-06-01 12:04:00 |
| 2 | add_to_cart | 2024-06-01 12:05:00 |
| 2 | checkout | 2024-06-01 12:07:00 |
| 3 | product_view | 2024-06-01 12:03:00 |
| 3 | product_view | 2024-06-01 12:06:00 |
Step 1: Initializing a Retentioneering Eventstream
Retentioneering uses the Eventstream object as its main abstraction. Let's convert our DataFrame into an eventstream, making sure to specify the correct column names:
es = re.Eventstream(
raw_data=df,
user_col='user_id',
event_col='event',
timestamp_col='timestamp'
)
Step 2: Exploring and Visualizing Event Flows
Let’s see how users move through the various stages of the journey. The es.plot_graph() function visualizes the event flow as a directed graph, where nodes represent events and edges represent transitions.
es.plot_graph(
weight_col='user_id', # Edge thickness shows number of unique users
layout='spring', # Graph layout
show_labels=True,
figsize=(10, 6)
)
This produces a flow graph, showing the most common paths users take. In a real notebook or script, you would see a network diagram where, for example, product_view leads to add_to_cart and so on.
Step 3: Building and Analyzing Conversion Funnels
A funnel is a sequence of steps you expect users to complete. In our e-commerce example, a typical funnel might be: product_view → add_to_cart → checkout → purchase.
Retentioneering makes funnel analysis simple:
funnel_steps = ['product_view', 'add_to_cart', 'checkout', 'purchase']
funnel = es.funnel(steps=funnel_steps)
# Display funnel statistics
funnel.stats()
The output will look like:
| Step | Users | Conversion (%) |
|---|---|---|
| product_view | 5 | 100% |
| add_to_cart | 5 | 100% |
| checkout | 3 | 60% |
| purchase | 2 | 40% |
This table shows how many users reach each step and the conversion rate from start to that step. (Numbers will vary with actual data.) You can also visualize the funnel:
funnel.plot()
Step 4: Retention and Cohort Analysis
Retention analysis helps you measure how many users return or perform a key action over time after their first activity. Cohort analysis groups users based on when they started and tracks their behavior.
ret = es.retention_matrix(
event='purchase', # Track retention based on purchase event
period='D', # 'D' for daily, 'W' for weekly, etc.
n_periods=7 # Days to track retention
)
ret.plot()
This generates a retention heatmap, where each row represents a cohort of users who first purchased on a given day, and each column shows the fraction returning in subsequent periods.
Retention Curve Math
Retention at period \( t \) is calculated:
\( R_t = \frac{N_t}{N_0} \)
Where:
- \( N_0 \) is the number of users in the cohort at period 0 (first activity)
- \( N_t \) is the number who returned or performed the target event at period \( t \)
Step 5: Path Analysis and Drop-off Exploration
Path analysis reveals the most common (or problematic) user journeys. Retentioneering can show frequent paths or transitions, helping you spot unexpected behaviors or drop-off points.
Top Paths
# Show top 5 most frequent event paths of up to 4 steps
paths = es.paths(max_len=4, top_k=5)
paths.head()
| Path | User Count |
|---|---|
| product_view → add_to_cart → checkout → purchase | 2 |
| product_view → product_view → add_to_cart → checkout | 1 |
| product_view → add_to_cart → checkout | 1 |
| product_view → product_view → add_to_cart | 1 |
Notice how some users repeat product_view before proceeding, or drop off before reaching purchase.
Drop-off Analysis
# Where do users most frequently drop out of the funnel?
funnel.dropoff()
This identifies the step with the highest loss of users, guiding UX improvements or targeted interventions.
Step 6: Segmenting and Comparing User Groups
Segmentation lets you compare behaviors between different user types. For instance, compare users who completed a purchase with those who didn't.
# Tag users who purchased
df['purchased'] = df['user_id'].isin(
df[df['event'] == 'purchase']['user_id']
)
# Split eventstreams by purchase status
purchasers = es.query_users(df[df['purchased'] == True]['user_id'])
non_purchasers = es.query_users(df[df['purchased'] == False]['user_id'])
# Compare their flow graphs
purchasers.plot_graph(title='Purchasers')
non_purchasers.plot_graph(title='Non-purchasers')
This comparison highlights distinct patterns: do non-purchasers get stuck at certain steps, or loop between views and carts? Such insights support targeted re-engagement or onboarding efforts.
Step 7: Advanced Usage – Sessionization and Custom Event Mappings
In real-world data, you may want to divide long user streams into sessions (periods of activity separated by inactivity) or map many raw events into broader categories.
Sessionization
# Split eventstream into sessions: gap > 30 min starts new session
es_sess = es.split_sessions(timeout=timedelta(minutes=30))
Custom Event Mapping
# Example: Map all product-related events to 'product_interaction'
event_map = {
'product_view': 'product_interaction',
'add_to_cart': 'cart_action',
'checkout': 'checkout',
'purchase': 'purchase'
}
es_mapped = es.map_events(event_map)
Step 8: Exporting Insights and Reporting
Once your analysis is complete, you can export tables or graphs for further reporting or dashboarding:
# Export funnel data to CSV
funnel.stats().to_csv('funnel_stats.csv', index=False)
# Save flow graph as image
fig = es.plot_graph(show=False) # Don't display
fig.savefig('user_flow_graph.png')
Case Study: Diagnosing Funnel Drop-offs
-
Suppose your analysis reveals a significant drop-off between
add_to_cartandcheckout. How can you use Retentioneering to diagnose and address this?- Visualize the flow graph with edge weights to spot alternative paths (e.g., do users return to
product_viewor exit after adding to cart?). - Segment users who abandon at this step and analyze their behavior: Do they visit certain categories or come from specific channels?
- Check the time users spend between
add_to_cartandcheckout. Delays might indicate friction (e.g., unexpected shipping costs, required account creation). - Test funnel changes or interventions, then measure their impact on conversion rates using the same funnel functions.
Integrating Retentioneering with Your Data Pipeline
You can integrate Retentioneering into your data pipeline by:
- Loading raw event logs directly from your data warehouse (e.g., via SQL or CSV export).
- Preprocessing and mapping events using pandas before creating your
Eventstream. - Automating regular analyses (e.g., daily funnel reports, retention curves) with scheduled scripts or notebooks.
# Example: Load events from a CSV exported from your data warehouse df = pd.read_csv('user_events.csv', parse_dates=['timestamp']) es = re.Eventstream(raw_data=df, user_col='user_id', event_col='event', timestamp_col='timestamp')
Summary Table: Key Retentioneering Functions
Function Purpose Example Usage Eventstream()Create eventstream from DataFrame es = re.Eventstream(raw_data=df, ...)plot_graph()Visualize event transitions as a graph es.plot_graph()funnel()Create and analyze conversion funnels funnel = es.funnel(steps=[...])retention_matrix()Generate retention cohort matrix es.retention_matrix(event='purchase')paths()Show top event paths es.paths(max_len=4, top_k=5)split_sessions()Break eventstreams into sessions es.split_sessions(timeout=...)map_events()Remap raw events to custom categories es.map_events(event_map)query_users()Filter eventstream by custom user lists es.query_users(user_ids)
Practical Applications in E-Commerce
Here are some real-world questions you can answer with Retentioneering:
- What percentage of users who view a product ultimately purchase?
- Where in the checkout flow do most users abandon their carts?
- Do users who come from paid ads behave differently than organic users?
- How does time to purchase vary by product category or device?
- Are returning users more likely to complete purchases than first-time visitors?
These insights can directly inform product improvements, marketing campaigns, and personalized user experiences.
Additional Resources
- Retentioneering Official Documentation
- Retentioneering on GitHub
- Pandas Library
- Seaborn Visualization Library
- Matplotlib Documentation
Ready to start? Install Retentioneering, load your data, and unlock the full story behind your user journeys today.
- Visualize the flow graph with edge weights to spot alternative paths (e.g., do users return to
Related Articles
- ANOVA Assumptions: Key Factors for Accurate Statistical Analysis
- Feature Engineering Techniques for Tabular Data in Machine Learning
- Python Slicing Examples for Data Analysis Explained
- Best Python Libraries for Beginners in Quantitative Analysis
- Feature Selection in Machine Learning: Methods and Best Practices