Time series analysis is the study of observations recorded in chronological order, like daily sales, hourly electricity demand or monthly hospital visits, so you can find trend, seasonality and autocorrelation and then use those patterns to forecast what happens next. To run one, plot the series, test whether it is stationary, transform it if it is not, fit a simple baseline before anything clever, and score the model on data held back in time order.
That order matters. Time series data breaks the independent-and-identically-distributed assumption that ordinary regression relies on, so methods that ignore temporal order give you tidy numbers that fail the moment you forecast. If you are wondering how to run a time series analysis for beginners without drowning in theory, the workflow below takes an afternoon and needs nothing more than a spreadsheet or a short Python script.
Table of Contents
- 1What You Need
- 2Real examples of time series data
- 3The five concepts you actually need
- 4Which software to use
- 5Step-by-Step: How to Run Time Series Analysis for Beginners
- 61. Define the Question and Prepare the Data
- 72. Explore the Series with a Line Plot
- 83. Check Whether the Series Is Suitable for Analysis
- 94. Transform the Data When Needed
- 105. Choose and Fit a Beginner-Friendly Model
- 116. Evaluate the Model and Forecast
- 127. Report the Analysis Clearly
- 13Common Mistakes
- 14Frequently Asked Questions
- 15How many time points do I need to run a time series analysis?
- 16Can I use time series analysis for a short or irregular series?
- 17Which software is best for a beginner learning time series?
- 18What does stationarity mean in time series analysis?
- 19How do I know whether my data has seasonality?
- 20How should I report a time series analysis in a research paper?
- 21Conclusion
What You Need
You need four things: a time variable, an outcome measured on a known schedule, a question with a forecast horizon attached to it, and one piece of software. That is the whole requirement list.
Real examples of time series data
Time series data is any set of measurements taken repeatedly over time on a fixed or near-fixed schedule. Common examples include:
- Daily share prices and trading volumes
- Monthly retail sales or monthly card spend
- Hourly website traffic or server load
- Daily maximum temperature at one weather station
- Quarterly GDP or monthly inflation readings
- Weekly hospital admissions by department
- Sensor readings from a water meter or a machine
Two features separate time series data from ordinary cross-sectional data. The order of rows matters, and the observations carry their own memory. A Monday figure is not an independent new draw, because it is conditioned on Sunday and Saturday, which is why a normal regression fit on the same numbers will usually look good in-sample and disappoint you out-of-sample.
The five concepts you actually need
Trend is the slow direction of the series. Seasonality is a pattern that repeats on a fixed cycle, weekly or yearly. Noise or residuals is what is left once trend and seasonality are removed. Stationarity means the mean and variance stay roughly constant over time, which is the condition most forecasting models assume. Autocorrelation is the correlation of the series with a shifted copy of itself, measured at each lag.
If those five are still fuzzy, keep reading. They come back in every step below.
Which software to use
Python and R give you the full workflow and proper forecast diagnostics. Excel is enough for plotting, decomposition and a basic trend projection, but it cannot test stationarity or fit ARIMA. SPSS and Stata handle panel and econometric work with time fixed effects and lagged dependent variables. SAS is the enterprise option when your company already runs it.
| Tool | Best fit | Weak spot |
|---|---|---|
| Python (pandas, statsmodels) | Full workflow from raw CSV to forecast with prediction intervals | Learning curve if you have never used a terminal |
| R (forecast, tseries) | Fastest route to reproducible forecasting and seasonal decomposition | Less familiar syntax outside academic settings |
| Excel | Quick plots, moving averages, trendlines, teaching yourself the ideas | No stationarity tests or ARIMA without VBA |
| SPSS | Courses that teach time series in a GUI, lagged variables | Forecast output is limited to a few methods |
| Stata | Econometric models with panel data and robust standard errors | Command syntax rather than click paths |
| SAS | Large institutional datasets and automated reporting | Heavy licensing and setup |
If you have never written code, read the steps below and then reproduce them in Excel. The logic is identical; only the menu names change.
Step-by-Step: How to Run Time Series Analysis for Beginners
There are three separate jobs hiding in the phrase time series analysis, and beginners mix them up. Exploration describes what the data does: plotting, decomposition, looking at the autocorrelation function. Forecasting produces future values. Formal model testing asks whether the assumptions of a statistical model hold. Do the exploration first, then the forecasting, then the testing, and never the other way round.
1. Define the Question and Prepare the Data
Write down the time variable, the outcome, the sampling frequency and how far ahead you want to predict. “Forecast weekly clinic visits for the next six weeks” is a usable question. “Analyse the clinic data” is not, because it hides the horizon that decides your model order and your validation split.
Then clean the frame. Sort rows chronologically, set the date column as the time index, and check that consecutive rows are actually one frequency apart. Unequal gaps quietly break models that assume equal spacing. Look for duplicate timestamps and decide whether to aggregate them to a sum or a mean. Note every gap and every missing value in a log file as you go, because you will need to explain those decisions later.
In pandas this is a few lines, and the same three checks apply in any tool:
import pandas as pd
df = pd.read_csv("visits.csv", parse_dates=["date"])
df = df.sort_values("date").set_index("date")
gaps = df.index.to_series().diff()
print(gaps.value_counts().head()) # are intervals equal?
print(df["visits"].isna().sum()) # how many missing values?
How you run a time series analysis for beginners is decided in this first step. Everything after it assumes your rows are in order and your frequency is real.
2. Explore the Series with a Line Plot
Plot the raw series against time before you compute anything. The plot is the highest-value diagnostic you have, and it takes ten seconds.
Look for four things. Direction tells you whether the level is rising or falling. Repeating waves at a fixed spacing tell you the seasonal period, so monthly data with a yearly hump means a period of 12. Sudden permanent jumps are level shifts, which usually mean a policy change or a recording change rather than a trend you can extrapolate. Isolated spikes far from their neighbours are outliers, and you should ask whether they are data errors before you smooth them away.
Gaps and ragged edges show up here too. A missing run in the middle of a chart explains a lot of odd behaviour later, so treat filling it as an analytical decision and record what you did.
3. Check Whether the Series Is Suitable for Analysis
A series is stationary when its mean, variance and autocorrelation structure stay roughly constant over time. Most ARIMA and exponential smoothing models assume a stationary series after transformation, and this is the step beginners skip.
Two tests cover it, and they test opposite null hypotheses:
| Test | Null hypothesis | Rejecting the null means |
|---|---|---|
| Augmented Dickey-Fuller | The series has a unit root, so it is non-stationary | The series is stationary |
| KPSS | The series is stationary around a trend | The series is non-stationary |
The ADF and KPSS disagreement is the single most common beginner question on econometrics and statistics forums. If ADF says stationary and KPSS says non-stationary, trust KPSS and apply differencing, because ADF has low power and often fails to detect a unit root in long series. Plot the series and a rolling mean together and decide visually too.
from statsmodels.tsa.stattools import adfuller, kpss
adf_p = adfuller(df["visits"])[1]
kpss_p = kpss(df["visits"], regression="c")[1]
print("ADF p:", adf_p, "| KPSS p:", kpss_p)
Read a p-value above 0.05 as no evidence against the null. Also plot the decomposition. STL breaks the series into a trend component, a seasonal component and residuals, which shows you at a glance whether seasonality is strong enough to model and whether the seasonal pattern is stable across years. Then look at the autocorrelation function and its partial counterpart, plotted as bars against lag. A slow decay across many lags means trend or non-stationarity. A clean spike at the seasonal lag confirms seasonality. On the PACF, a spike that dies after the first few lags suggests a low autoregressive order.
4. Transform the Data When Needed
If the series is not stationary, fix it before modelling, and change one thing at a time so you know what worked.
- Log transformation tames variance that grows with the level, which suits sales counts and demand figures.
- Detrending removes a fitted straight line and suits data whose growth you do not want to project.
- First differencing removes a stochastic trend, and it fixes most series that drift upward.
- Seasonal differencing subtracts the value one season earlier, which is step 12 for monthly data with a yearly cycle.
- Resampling changes the frequency, so daily totals become weekly or monthly figures.
After each change, re-run the test and re-plot. Stop as soon as the series is stationary. Over-differencing is a real failure mode: it flattens real structure and inflates the forecast error. Write down the exact order of every transformation in order, because you must reverse them to report forecasts in the units your reader understands.
y = np.log(df["visits"])
y = y.diff().dropna() # first difference
y = y.diff(12).dropna() # seasonal difference for monthly data
5. Choose and Fit a Beginner-Friendly Model
Start with the cheapest model that could work, because a complicated model that cannot beat a naïve forecast has not earned its place.
The naïve forecast predicts that the next value equals the last observed value. The seasonal naïve predicts that next year equals this year. Both take a second to compute and they are the benchmark everything else must beat. A moving average smooths the last few values and is fine when you only need a stable baseline. Exponential smoothing, in its simple and Holt-Winters forms, weights recent observations more heavily and handles level and seasonal structure. A regression on time plus seasonal dummies works well when you want an explainable trend. ARIMA, written as p, d, q, models an autoregressive part, a differencing part and a moving average part, and SARIMA adds a seasonal block.
Pick p and q from the plots rather than by habit. If the ACF tails off gradually and the PACF cuts off after lag p, choose a small autoregressive order such as an ARIMA with one autoregressive term. If the ACF cuts off and the PACF tails off, favour the moving average side. Use the AIC from the fitted candidates as a tiebreaker, and let differencing tell you d instead of choosing it by feel.
from statsmodels.tsa.arima.model import ARIMA
fit = ARIMA(train, order=(1,1,1), seasonal_order=(0,1,1,12)).fit()
print(fit.summary())
6. Evaluate the Model and Forecast
Split by time, never by chance. Shuffling rows before splitting leaks future values into training and produces an accuracy figure you cannot reproduce in production. Hold out the final block of your series as the test set, fit on everything before it, and forecast across the held-out block.
Rolling-origin cross-validation does the same idea repeatedly: train up to an earlier cut-off, forecast forward, then move the cut-off later and repeat. It tells you whether the model held up across different periods instead of one lucky window. Compare your candidates on the same folds and choose the winner on error, not on how complicated it looks.
| Metric | What it measures | Use it when |
|---|---|---|
| MAE | Average size of the miss in the original units | You need an error a non-technical reader can picture |
| RMSE | Misses with large errors punished harder | Big misses are much more costly than small ones |
| MAPE | Average miss as a percentage of the actual value | Values are never near zero and you need a relative figure |
Compare the error against the naïve baseline, because a five percent MAPE means nothing if last month’s guess already hit seven percent. Then reverse every transformation you applied, exponentiating log values and summing differenced ones back up, and produce a prediction interval rather than a single number. Anyone forecasting for a decision needs the range, not the point. If your model beats nothing on held-out data, go back to step 2 rather than adding more parameters.
7. Report the Analysis Clearly
A readable write-up covers eight things in this order: the question and horizon, the data structure and frequency, what the plot showed, which transformations you applied and why, which model you chose and what beat it, how you validated, the forecast with its prediction interval, and the limits.
Limits deserve a sentence of their own. Short series, an unstable seasonal pattern and a level shift inside the training window all weaken any forecast you produce. Finish with plain language: next quarter’s visits are likely to fall between 900 and 1,050, centred near 970, and that range assumes the staffing pattern holds. If you are unsure how to run a time series analysis for beginners in a report your manager will read, that sentence is the deliverable.
Common Mistakes
Most failed beginner analyses come from a short list of avoidable choices.
Treating dates as ordinary numbers. A regression on a row index or a raw timestamp implies a linear gap you did not observe, and it throws away the lag structure. Fix: use the date as the index, work in frequency steps, and derive lag features explicitly.
Shuffling or randomly splitting the data. Future observations leak into training and every accuracy number inflates. Fix: chronological split, then rolling-origin cross-validation for the model comparison.
Fitting ARIMA before a naïve baseline. You end up defending a complicated model that a one-line guess already matched. Fix: always compute the naïve and seasonal naïve error first, and keep only models that beat it.
Over-differencing. Differencing until the p-value looks good destroys real structure and makes forecasts worse. Fix: difference once, retest, and stop as soon as the series is stationary.
Ignoring seasonality. Monthly retail data with a December peak needs seasonal differencing or Holt-Winters. Fix: run STL and look at the seasonal panel before choosing the model.
Believing the wrong stationarity test. ADF rarely rejects in long trending series. Fix: when ADF and KPSS disagree, trust KPSS and difference the series.
Reporting forecasts in transformed units. Readers get log values or differenced numbers and assume they are visits or dollars. Fix: write down the transformation order and reverse it before reporting.
Publishing a single forecast number. One number invites false certainty. Fix: always attach prediction intervals and say how they were produced.
Reading autocorrelation as causation. Two series moving together do not cause each other, and common trends create spurious correlation. Fix: difference both, test for causality with an explicit model, and keep the language descriptive.
Before you submit, confirm the rows are in chronological order, the plot came first, stationarity was tested and fixed, a baseline was beaten, the split respected time, transformations were reversed, and an uncertainty range accompanies every forecast.
Frequently Asked Questions
How many time points do I need to run a time series analysis?
For a monthly series, most beginners need at least 36 observations, and 60 or more is better, because a yearly seasonal cycle needs three full cycles before any pattern is trustworthy. Daily data gives you far more points but often less usable history. As a rule, aim for at least four seasonal cycles for any method that models seasonality. Below that, describe the series with a plot and a moving average and skip formal forecasting.
Can I use time series analysis for a short or irregular series?
Short series can still be explored. Plot it, run a moving average, and report the pattern qualitatively. For irregular spacing, resample to a common frequency first, because most models assume equal intervals, and state clearly how you filled the gaps. If your observations are genuinely sporadic rather than evenly spaced, look at state space models, which handle uneven timing better than ARIMA.
Which software is best for a beginner learning time series?
Excel is the easiest place to learn the ideas, since plots, moving averages and trendlines take minutes and the concepts transfer everywhere. Python with pandas and statsmodels is the better long-term choice because it handles stationarity tests, decomposition and forecast intervals in one place. R sits between the two. Pick the tool your workplace or course already uses, then move on to the method rather than the software.
What does stationarity mean in time series analysis?
A stationary series has a roughly constant mean, variance and autocorrelation structure across time, so its behaviour today resembles its behaviour yesterday. Most ARIMA and exponential smoothing models assume this after transformation. Test for it with the augmented Dickey-Fuller and KPSS tests, then fix violations with a log transform, detrending or differencing until the tests agree.
How do I know whether my data has seasonality?
Run an STL decomposition and look at the seasonal panel. A repeating shape of similar height across cycles means seasonality is stable. You can also check the autocorrelation function for a clear spike at the seasonal lag, such as 12 for monthly data or 7 for daily data. If the seasonal panel grows or shrinks over time, the pattern is unstable and a seasonal model will struggle.
How should I report a time series analysis in a research paper?
Report the question and horizon, the data frequency and length, the transformations you applied, the stationarity test results, the candidate models you compared, your validation method, and the forecast with its prediction interval. Give the error metric against the naïve baseline so readers can judge whether the model earned its complexity. Finish with a plain-language statement of what the series is likely to do next and the main limitations.
Conclusion
Start by loading your dated data and plotting it. Everything else in how to run a time series analysis for beginners follows from what that first chart shows you, and skipping it is why so many first models fail.
From there the sequence is fixed. Check the structure with decomposition and the autocorrelation plots, test for stationarity and fix what fails, compare a baseline against a couple of candidates on a chronological split, reverse your transformations, and report a range instead of a single number. Work through those seven steps on your own dataset before you look at anyone else’s tutorial, because the plot will always surprise you.


