Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
espe
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
espe
3y ago
thanks for the clarification. if your base population is that large then it's frequencies and you get a fingerprint. well done.
32.
▲
by
espe
3y ago
very efficient but also brittle. that must be vast amounts of relatively clean data. you have to magically set the number of top n words to in- and exclude. for most user generated content one would need to heavily normalize the text, e.g.
33.
▲
by
espe
3y ago
+1 for setfit. a baseline that's hard to beat.
34.
▲
by
espe
3y ago
it's all sshfs.. afaik they're looking into changing that. i don't have the i/o issues but always backup any files in the vm as they tend to get lost at some point.
35.
▲
by
espe
3y ago
having worked with them for a few years, i'd say it's one of the most innovative library institutions and we (in the german speaking world and beyond) need them. they produced countless digital corpora and editions. largely histor
36.
▲
by
espe
3y ago
lima is nice. just beware that os updates can accidentially nuke the vm. got to try out utm sometime.
37.
▲
by
espe
3y ago
mostly agree. not only a notation tool though - you can see some familiarity to scheme, that data is also an actionable description, and literate programming, that this actionable description is also a human readable description. together w
38.
▲
by
espe
3y ago
it is easy to paint XML as anachronistic, and as crude serialization format it certainly is. but as many analogies, that only goes so far. there is still the excellent tooling (XSLT, DTD, schematron..) and the fact that there are elaborate
39.
▲
by
espe
3y ago
regarding the relationship: yes, and in most ways it probably is a subset. is there really such a set of rules that generate all possible sentences? in any case i wanted to say the materiality and cultural activity heavily influences what c
40.
▲
by
espe
3y ago
nicely put! many aspects of text at least historically have much to do with its materiality (also in a cognitive development sense, learning how to write etc.). what we can think about nowadays is that text and speech might not be a necessa
41.
▲
by
espe
3y ago
text and language intersect. in some ways, text is a superset of language, mostly due to social, or what is also called pragmatic, factors that complement semantics. also, the semantics/syntax interface is everything else than clear cu
42.
▲
by
espe
3y ago
compression :)
43.
▲
by
espe
3y ago
from a linguistic standpoint a text is a whole lot more than language: it is an externalisation of thought that is fixed onto a medium using writing utensils and most of all, cultural norms in the form of a wild variety of different genres
44.
▲
by
espe
4y ago
actually that might not be the case. don't underestimate the value of older, better understood and much smaller models. also, why not call bert-style (encoder) models LLMs as well. i would expect last-gen models to give us an edge in c
45.
▲
by
espe
4y ago
you can think of fine-tuning as rewiring where it matters/can be probed. a kind of exhaustive reorganisation of the latent model space given some seed statement like you describe might be possible with LLMs that are jointly trained as
46.
▲
by
espe
4y ago
exactly. visual information is more compressable than natural language: much of it boils down to locality, whereas language forms are highly pareto distributed plus the conceptual system is a huge hypergraph, so it's rather the opposit
47.
▲
by
espe
4y ago
i fail to see how this data cleaning could not be solved with proper tokenization and some distance measure. the amount of power used for those api calls is slighty obscene. edit: don't want to rant. it's not a bad post and i'
48.
▲
Automatically Detecting Offensive Language in the German Speaking Twittersphere
(blog.hexis.ai)
1 points
by
espe
7y ago
|
0 comments
49.
▲
Show HN: Hate Speech and Abuse Detection in Text-Based Communication
(hexis.ai)
1 points
by
espe
7y ago
|
0 comments
50.
▲
by
espe
7y ago
last minute typo correction.. sorry for that!
51.
▲
Show HN: Text-Based Hate Speech and Abuse Detection That Works
(hexis.ai)
1 points
by
espe
7y ago
|
2 comments
52.
▲
Offensive Language in the German Speaking Twittersphere
(blog.hexis.ai)
1 points
by
espe
7y ago
|
0 comments
53.
▲
by
espe
9y ago
it would ultimately more fruitful to treat irony as a semantic phenomenon, as belonging to the spectrum of tropes such as metaphor, metonymy and synecdoche. how should treating it as a violation tell you more about how it works?
54.
▲
by
espe
9y ago
claims too much. and the task definition seems off: sarcasm is more like a speaker attitude, while irony is the linguistic phenomenon. another case of "lets throw an ANN onto anything".
55.
▲
by
espe
11y ago
afaik, stylo ( https://sites.google.com/site/computationalstylistics/stylo ) is the academic go-to solution. it is even sporting a gui