When two people sing a song together, the vocals stem holds both of them. There is no way to hear one singer on their own, so anything that depends on a single voice is out of reach: learning one part, muting one singer, checking your own pitch against one performance.
The lead/backing vocal split (#275) does not answer this. It pulls a lead away from the harmonies sitting behind it, which is a different question from two people trading verses.
What makes this hard:
- Two singers in different vocal ranges and two singers in the same range are separate problems. A model that is good at one is not good at the other.
- Whether a split actually happened cannot be measured from the audio without already knowing the answer. A wrong separation is confident and uncorrelated in the same way a right one is, so the usual blind quality checks will pass a split that never occurred.
- Once a full chorus joins, a model with two outputs has nowhere to put a third voice.
Constraints worth knowing before starting: there is no training budget here, so this has to be built from published weights, and the licence on those weights has to allow redistribution.
When two people sing a song together, the vocals stem holds both of them. There is no way to hear one singer on their own, so anything that depends on a single voice is out of reach: learning one part, muting one singer, checking your own pitch against one performance.
The lead/backing vocal split (#275) does not answer this. It pulls a lead away from the harmonies sitting behind it, which is a different question from two people trading verses.
What makes this hard:
Constraints worth knowing before starting: there is no training budget here, so this has to be built from published weights, and the licence on those weights has to allow redistribution.