4 ms·
u may want to use regex to reduce / eliminate if checkings i use newlisp, this is the code: (set 'urls '( "http://motors.shop.ebay.com/Parts-Accessories_Car-
by hs 18y ago
u may want to use regex to reduce / eliminate if checkings
i use newlisp, this is the code:
(set 'urls '(
"http://motors.shop.ebay.com/Parts-Accessories_Car-Truck-Parts-Accessories__f350-wheels-20_W0QQ_fxdZ1QQ_osacatZPartsQ2dAccessoriesQQ_trksidZm270Q2el1313&caz.html http://motors.shop.ebay.com/Parts-Accessories_Car-Truck-Part..."
"http://cgi.ebay.com/ebaymotors/FACTORY-15-Mercedes-E320-300E-OEM-Chrome-Wheels-Rims_W0QQitemZ290277158739QQihZ019QQcategoryZ43955QQssPageNameZWDVWQQrdZ1QQcmdZViewItem&caz.html http://cgi.ebay.com/ebaymotors/FACTORY-15-Mercedes-E320-300E..."
"http://shop.ebay.com/items/__mini-cooper-rims?_trkparms=72%3A543%7C66%3A2%7C65%3A12%7C39%3A1&caz.html http://shop.ebay.com/items/__mini-cooper-rims?_trkparms=72%3..."))
(define (getKW url)
(find {([^/|^_]*)(_W0QQ|\?)} url 1) ;find using regex
$1) ;return the first matched string inside (bla*)
(map println (map getKW urls))
;f350-wheels-20
;FACTORY-15-Mercedes-E320-300E-OEM-Chrome-Wheels-Rims
;mini-cooper-rims
- matth 18y agoThe particular hangup I faced when using regex was with this URL: http://vi.ebaydesc.com/ebaymotors/ws/eBayISAPI.dll?ViewItemDesc&item=290277158797&t=1227377775000&ds=2&seller=la-wheel-and-tire&js=e583:1&hr=http://shop.ebay.com/?_from=R40&caz.html http://vi.ebaydesc.com/ebaymotors/ws/eBayISAPI.dll?ViewItemD... >>> x = re.compile('(?:(item|t|hr)=(\d+))') >>> x.findall(url) [('item', '290277158797'), ('t', '1227377775000')] I can't for the life of me figure out how to get the hr value. I started messing around with regex again, but no luck so far. I can do the same with Python's string methods: >>> hr_start = url.find('hr=') >>> hr_end = url.find('&',url.find('hr='),len(url)) >>> url[hr_start:hr_end] 'hr=http://shop.ebay.com/?_from=R40' http://shop.ebay.com/?_from=R40' It's just messy as hell, it'd be nice to do everything in one swoop.
- matth 18y agoOk, mission accomplished methinks: >>> x = re.compile("[\\?&](seller|item|hr)=([^&#]*)") >>> x.findall(url) [('item', '290277158797'), ('seller', 'la-wheel-and-tire'), ('hr', 'http://shop.ebay.com/?_from=R40' http://shop.ebay.com/?_from=R40')]