Brainsuite follows a strict "science-first" strategy when developing AI pipelines. For video apps these means that we leverage a state-of-the-art scene detection model to segment videos into meaningful segments ("scenes"). Why? Because this is exactly what the human brain does - instead of processing each frame seperately, the human brain leverages what is called "event segmentation" where a temporal event (like a video) is split into segments (e.g. scenes).
Within each scene of a video, our video AI pipeline extracts so-called key frames - frames within the scene that are most representative of what is shown. These key frames then form the basis of the scene and video KPIs. We are excited to the launch of an updated key frame selection approach. The main focus is to select even more accurately all frames in a video that show brand and product. To this end our experts manually labelled 50 videos with detailled information about when brand and/or product are shown.
In an exhaustive analysis leveraging different optimization strategies we compared our previous key frame selection method against a large number of alternatives. This effort resulted in an improvement of brand/product recall from previously 92% (i.e. of 100 brand/product frames, 92% were selected) to more than 96%. This is the solution that is now live on Brainsuite. We are committed towards continously reviewing and improving all our KPIs towards even more accuracy.