Bless your heart OP, but attempting to hierarchically taxonomise all of the internet is just futile.
And not just even talking about the infrastructure required, but the fact that pinning ideas into a taxonomy or ontology frames them in image of the taxonomising taxonomer, and not in the one querying it.
It is also one of the reasons that the ideas of Web 2.0 never truly caught on, and why the population of json users dwarfs that of RDF users.
Perhaps instead of a hierarchical categorisation you could use vector embeddings with a similarity measure. That frees you from the burden of taxonomising and perhaps leaves you more energy to refine your indexed corpus for relevance and utility.
Bless your heart OP, but attempting to hierarchically taxonomise all of the internet is just futile.
And not just even talking about the infrastructure required, but the fact that pinning ideas into a taxonomy or ontology frames them in image of the taxonomising taxonomer, and not in the one querying it.
It is also one of the reasons that the ideas of Web 2.0 never truly caught on, and why the population of json users dwarfs that of RDF users.
Perhaps instead of a hierarchical categorisation you could use vector embeddings with a similarity measure. That frees you from the burden of taxonomising and perhaps leaves you more energy to refine your indexed corpus for relevance and utility.
Best of luck!