4 ms·
You don't really often need an array language, just like you don't really often need regexes. When when you have a problem that perfectly fits the bill, they a
by BiteCode_dev 10mo ago
You don't really often need an array language, just like you don't really often need regexes.
When when you have a problem that perfectly fits the bill, they are very good at it.
The problem is they are terrible at everything else. I/O, data validation, manipulation strings, parsing, complex logic trees...
So I feel like just like regexes, there should be an array language parser embedded in most languages, that you opt in locally for just this little nudge.
In Python, it would be nice to be able to "import j" like you "import re" in the sdlib.
The entire J code base, including utility scripts, a console, a stdlib and a regex engine, is 3mb.
- ogogmad 10mo agoI suspect the same regarding the analogy with regex, but I still haven't finished learning an array language. Do you know what you'd use an array language for?
- BiteCode_dev 10mo agoPersonally, to defer importing numpy until I can't anymore. Sometimes you just need a little matrix shenanigans, and it's a shame to have to bring in a bazooka to get decent ergonomics and performances.
- ogogmad 10mo agoI don't know if you're aware that there's a formal analogy between matrix operations and regex operations: Matrix vs Regex -------------- A+B with A|B A*B with AB (1 - A)^{-1} with M* To make the analogy between array programming and regex even more precise: I think you might even be able to make a regex engine that uses one boolean matrix for each character. For example, if you use the ASCII character set, you'd use 127 of these boolean matrices. The matrices should encode transitions between NFA states. The set of entry states should be indicated by an additional boolean vector; and the accepting states should be indicated by one more boolean vector. The regex operations would take 1 or 2 NFAs as input, and output a new NFA.
- BiteCode_dev 10mo agoDidn't know that but I assume you can share most of the engine's logic anyway. Those kind of generalisations tend to break down once you get pratical implementations.
- ogogmad 10mo agoThe following is a Python prototype: import numpy as np from scipy.sparse import csr_matrix, bmat class NFA: def __init__(self, T, S, E): self.T = T; self.S = S; self.E = E @property def null(self): # Nullable? return (self.S.T @ self.E)[0,0] # --- 1. The Core Algebra --- def disjoint(fa1, fa2): """ Places fa1 and fa2 into a shared, non-overlapping state space. """ n1, n2 = fa1.S.shape[0], fa2.S.shape[0] z = lambda r, c: csr_matrix((r, c), dtype=bool) # Block Diag Transitions chars = set(fa1.T) | set(fa2.T) T_new = {} for c in chars: m1 = fa1.T.get(c, z(n1, n1)) m2 = fa2.T.get(c, z(n2, n2)) T_new[c] = bmat([[m1, None], [None, m2]], format='csr') # Stack Vectors S1 = bmat([[fa1.S], [z(n2,1)]], format='csr') S2 = bmat([[z(n1,1)], [fa2.S]], format='csr') E1 = bmat([[fa1.E], [z(n2,1)]], format='csr') E2 = bmat([[z(n1,1)], [fa2.E]], format='csr') return NFA(T_new, S1, E1), NFA(T_new, S2, E2) def fork(fa): """ Returns two references to the exact same machine. """ return fa, fa def connect(fa1, fa2): """ The General "Sequence" Op. Wires fa1.End -> fa2.Start, and updates Start/End vectors. """ # 1. Transitions: T_new = T_combined + (Bridge @ T_combined) # Bridge = Start_2 * End_1^T Bridge = fa2.S @ fa1.E.T chars = set(fa1.T) | set(fa2.T) T_new = {} for c in chars: # If fa1==fa2 (fork), this just gets fa1.T[c] # If fa1!=fa2 (disjoint), this adds the non-overlapping blocks m_comb = fa1.T.get(c, _z(fa1)) + fa2.T.get(c, _z(fa2)) # Apply the feedback/feedforward T_new[c] = m_comb + (Bridge @ m_comb) # 2. States: Standard Concatenation Logic # S_new = S1 + (S2 if N1) # E_new = E2 + (E1 if N2) # Note: If fa1==fa2, this correctly computes S + (S if N) = S S_new = fa1.S + (fa2.S if fa1.null else _z(fa1, 1)) E_new = fa2.E + (fa1.E if fa2.null else _z(fa1, 1)) return NFA(T_new, S_new, E_new) # --- 2. The Operations (Now Trivial) --- def cat(fa1, fa2): return connect(*disjoint(fa1, fa2)) def leastonce(fa): return connect(*fork(fa)) def union(fa1, fa2): d1, d2 = disjoint(fa1, fa2) chars = set(d1.T) | set(d2.T) T_sum = {c: d1.T.get(c, _z(d1)) + d2.T.get(c, _z(d2)) for c in chars} return NFA(T_sum, d1.S + d2.S, d1.E + d2.E) def star(fa): return union(one(), leastonce(fa)) # --- Helpers --- def lit(char): T = {char: csr_matrix(([True], ([1], [0])), shape=(2,2), dtype=bool)} return NFA(T, _v(1,0), _v(0,1)) def one(): return NFA({}, _v(1), _v(1)) # Epsilon def _z(fa, c=None): return csr_matrix((fa.S.shape[0], c if c else fa.S.shape[0]), dtype=bool) def _v(*args): return csr_matrix(np.array(args, dtype=bool)[:, None]) # --- Execution --- def run(fa, string): curr = fa.S for char in string: if char not in fa.T: curr = _z(fa, 1) else: curr = fa.T[char] @ curr return (curr.T @ fa.E)[0,0] if __name__ == "__main__": # Test: (a|b)+ c # Logic: cat( leastonce( union(a,b) ), c ) a, b, c = lit('a'), lit('b'), lit('c') regex = cat(leastonce(union(a, b)), c) print(f"abac: {run(regex, 'abac')}") # True print(f"c: {run(regex, 'c')}") # False (Needs at least one a/b)
- shawn_w 10mo ago>... there should be an array language parser embedded in most languages, that you opt in locally for just this little nudge. April is this for Common Lisp: https://github.com/phantomics/april https://github.com/phantomics/april
- avmich 10mo agoAFAIK APL was used to verify the design of IBM 360, finding some flaws. I wrote my first parser generator in J. I think these both contradict your opinion. I think Iverson had a good idea growing language from math notation. Math operations often use one or few letters - log, sin, square, sign of sum or integral. Math is pretty generic, and I believe APL is generic as well.
- BiteCode_dev 10mo agoPeople wrote video games in Excel.