Adding rows in a dataframe for missing days for each entity

Question

Adding rows in a dataframe for missing days for each entity

66 Views Asked by Salocin R. At 06 June 2025 at 21:02

I have the following problem: my dataframes look something like this:

ID Date        Value

1 2016-06-12   2
1 2016-06-13   2.5
1 2016-06-16   4
2 2016-06-12   3
2 2016-06-15   1.5

As you can see I have missing days in my data. So I much rather want something like this:

ID Date        Value

1 2016-06-12   2
1 2016-06-13   2.5
1 2016-06-14   NaN
1 2016-06-15   NaN
1 2016-06-16   4
2 2016-06-12   3
2 2016-06-13   NaN
2 2016-06-14   NaN
2 2016-06-15   1.5

In order to solve that I did the following:

df_new = df.groupby('ID').apply(lambda x: x.set_index('Date').resample('1D').first())

This solution works, but takes about half an hour to process a large dataset. Thus, I wanted to know whether if is there a better solution?

Original Q&A

There are 1 best solutions below

**jezrael** · Answer 1

First idea is create all posible combinations of ID and Date values with and then merge with left join:

from  itertools import product

df['Date'] = pd.to_datetime(df['Date'])

L = list(product(df['ID'].unique(), pd.date_range(df['Date'].min(), df['Date'].max())))

df = pd.DataFrame(L, columns=['ID','Date']).merge(df, how='left')
print (df)
   ID       Date  Value
0   1 2016-06-12    2.0
1   1 2016-06-13    2.5
2   1 2016-06-14    NaN
3   1 2016-06-15    NaN
4   1 2016-06-16    4.0
5   2 2016-06-12    3.0
6   2 2016-06-13    NaN
7   2 2016-06-14    NaN
8   2 2016-06-15    1.5
9   2 2016-06-16    NaN

Or use DataFrame.reindex, but performance should be worse, depends of data:

df['Date'] = pd.to_datetime(df['Date'])

mux = pd.MultiIndex.from_product([df['ID'].unique(), 
                                  pd.date_range(df['Date'].min(), df['Date'].max())],
                                  names=['ID','Date'])

df = df.set_index(['ID','Date']).reindex(mux).reset_index()
print (df)
   ID       Date  Value
0   1 2016-06-12    2.0
1   1 2016-06-13    2.5
2   1 2016-06-14    NaN
3   1 2016-06-15    NaN
4   1 2016-06-16    4.0
5   2 2016-06-12    3.0
6   2 2016-06-13    NaN
7   2 2016-06-14    NaN
8   2 2016-06-15    1.5
9   2 2016-06-16    NaN

Adding rows in a dataframe for missing days for each entity

There are 1 best solutions below

Related Questions in PYTHON

Related Questions in PANDAS

Related Questions in DATAFRAME

Related Questions in MISSING-SYMBOLS

Trending Questions

Popular # Hahtags

Popular Questions